Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T17:37:39.675692Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 2 inbound Pith citation observations for arXiv:2502.00840.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T17:37:39.675692Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:00:19.630547Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-04T23:10:21.867778Z
92 of 92 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e2140333-e41d-4262-babf-344189700f7e · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Chat generative pre-trained transformer (chat- gpt)
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cea145c-5b3c-45de-a31e-bd36475a504c · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ba827a8-8fac-4239-962e-599a67040617 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Gemma: Open Models Based on Gemini Research and Technology
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74610e66-12fc-4343-8437-0fcfe916ef52 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Mistral 7B
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01e4ab5a-08e7-4396-9012-3dfd8245a74f · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense The Falcon Series of Open Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64401fd9-c23c-4d59-aa32-5ca14ea0d8f3 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Qwen Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0ec0fee-50b6-4ebe-92af-222ca99d8f52 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Gaussian Error Linear Units (GELUs)
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c704735d-6f10-4dcb-b2a8-67c736845040 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sigmoid- weighted linear units for neural network function ap- proximation in reinforcement learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6aaae0c-d6cd-4a15-bb84-41a7bbc6a0cb · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Awq: Activation- aware weight quantization for on-device llm compres- sion and acceleration
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35836eb3-c6d7-437f-8d81-84cfea072082 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense GPTQ: Accurate post-training compression for generative pretrained transformers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 479f9a15-8326-4792-8bfd-ec4a0f8371c4 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Billm: Pushing the limit of post-training quan- tization for llms
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 651d0dfb-4ab2-45ee-9c83-d4503e1b4c98 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sparsegpt: Massive lan- guage models can be accurately pruned in one-shot
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bbb1169-9b05-45db-bc95-7c1ca97e02af · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Llm- pruner: On the structural pruning of large language mod- els
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7a430ef-fed4-41b6-b1ba-6de630ba7c3e · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faed12e6-880d-4a75-9f5a-8bbbb02903de · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Structured Pruning of Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 058c99be-1b39-4bf5-9a2b-a218ac461f99 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Fast inference from transformers via speculative decoding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68d3aff6-70f9-42cc-b2b5-c48b44aca84b · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Speculative decoding with big little de- coder
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdda5e25-d4fe-4fb3-92b7-8fe140a74b12 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Iron: Private infer- ence on transformers
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5dafbbe-a6ca-40c3-a41c-8fb73b586f97 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Ciphergpt: Secure two- party gpt inference
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be552ba8-a521-4ba9-9f1b-57b70edcef59 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Bolt: Privacy-preserving, accu- rate and efficient inference for transformers
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19367cde-fe3b-4f0e-b017-bbd0653347c2 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense BumbleBee: Secure Two-party Inference Frame- work for Large Transformers
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1b0a96f9-7cf6-43b7-8d4a-bccdc494eec1 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Secure transformer infer- ence made non-interactive
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4e6321c5-dda7-4236-96aa-7c40970326c2 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Training-Free Activation Sparsity in Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 287bf1e2-4352-430b-ac76-5ab02d64ee48 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcc82b43-ba2b-4f9a-b7d9-6cbb855b197e · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Relu strikes back: Exploiting activation sparsity in large lan- guage models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 288148a7-7b86-42e1-af4f-39648d09b0d8 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Smoothquant: Accu- rate and efficient post-training quantization for large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a1d2126a-2aa5-4d7a-869f-5e4a87678d1a · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Omniquant: Omnidirection- ally calibrated quantization for large language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8c840334-b818-4007-91e1-3b4854e5963b · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sirnn: A math library for secure rnn inference
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4f1e32b9-8c20-42b2-84ef-e4f526db4df2 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense From chatgpt to threat- gpt: Impact of generative ai in cybersecurity and privacy
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 637b992f-b3f1-45cf-ba58-87c2879cdb14 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Exploiting large language models (llms) through decep- tion techniques and persuasion principles
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4e4e6dbc-4e9c-4b45-afbe-ce3af4e2d85b · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ef5ecec-76c2-49c2-9151-df6e15127c28 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Training language models to follow instructions with human feedback
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c76fa3c1-23e2-4af4-95c2-8bb3a4510ec2 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Safe RLHF: Safe Reinforcement Learning from Human Feedback
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 990fe8fb-ecd8-4a21-a561-8001f009b155 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Di- rect preference optimization: Your language model is secretly a reward model
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 28292dea-c199-4d8d-b0c6-068dbfc45920 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21dcb9ae-d6d4-444f-96b0-4f402b7abfbe · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 648e5c71-b1cd-46a7-b4ce-4f903832b0df · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 554ca386-87ae-4910-a4bb-96a9a97df6b0 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense A simple and effective pruning approach for large lan- guage models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation eb45a763-0ab2-44c5-be77-10d227b55ad4 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Dynamic context pruning for efficient and interpretable autoregressive transformers
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7ce3b523-7318-4acf-8037-686931756f86 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Adapting language models to compress contexts
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c805132b-81c1-4aa9-86dc-05ae00074791 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache manage- ment
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 71451ffd-5c83-4a1e-a185-f5420ad3677c · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense {IMPRESS}: An {Importance- Informed}{Multi-Tier} prefix {KV} storage system for large language model inference
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3a1c35c4-aa16-4794-9217-f105d484939a · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e7788879-bce4-4f23-b5e5-cc32f49dd2cb · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Kvquant: Towards 10 million con- text length llm inference with kv cache quantization
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 10aad630-7787-44c9-855c-ac005805c07b · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sequoia: Scalable and robust speculative de- coding
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 48c61911-7bba-476a-bea3-44a973b2df16 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Decoding compressed trust: Scrutinizing the trustworthiness of efficient llms under compression
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 45a28834-227a-4d5b-8201-4e64193306de · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Harmlevelbench: Evaluating harm-level compliance and the impact of quantization on model alignment
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5b8a9931-e89a-415e-892e-742fb4724cb2 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Exploiting llm quantization
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ceffec12-a271-4d19-b0cb-f420f35f503f · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 58b43f76-c1a2-4ce1-b0e4-34a321323825 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Qllm: Accurate and efficient low-bitwidth quantization for large language models.In- ternational Conference on Machine Learning, 2024
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4b9e6550-0094-4b81-ad21-6164bfb5beee · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Introducing meta llama 3: The most capable openly available llm to date
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 036314c3-df4a-428e-85ef-917363701540 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Mpcformer: fast, performant and private transformer inference with mpc
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bdb1cc3d-e089-4a1e-9e20-bf96c91f299b · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Encryption-Friendly LLM Architecture
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f2ac483-69c7-42ab-bbfb-31c9b162cb94 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Deja vu: Con- textual sparsity for efficient llms at inference time
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6da7ef52-b474-4671-86df-2cb90b6c996c · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Cats: Context-aware thresh- olding for sparsity in large language models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d9906681-c789-4184-9a25-d9900b8443bc · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Judging llm-as- a-judge with mt-bench and chatbot arena
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04a0dd48-3eb6-413b-94ef-63469149a2a4 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Assessing the brittleness of safety alignment via pruning and low-rank modifica- tions
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f9eaa58c-f081-48d1-9de3-5221b7b1162b · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b32f45a2-16e6-49de-8886-851d2c32f982 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Lisa: Lazy safety alignment for large language models against harmful fine-tuning at- tack
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a5fb86ba-5cd5-4115-adb7-184980feab14 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2012904-6ca2-4b52-8886-23136f88c45a · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e98b3c49-8494-4c98-9d5a-0b2b8460beac · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Detoxifying Large Language Models via Knowledge Editing
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ed5edc6-4186-4910-8e00-71df8537532c · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Recovering the Pre-Fine-Tuning Weights of Generative Models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dcecaae4-378e-4c13-a3a7-3cc5ef393661 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d6d80ec-20af-4486-bbc5-73332f878ca0 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1a338de-b2c4-4f79-8bcc-dacded386d85 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a15da23-7dd0-4138-8b6c-b2475b3ffc78 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Compressing LLMs: The Truth is Rarely Pure and Never Simple
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 960c1af9-ac41-4784-877c-0b3489ae88f3 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d45dd858-0691-4600-b9dd-4eca0d8c4165 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Discovering sparsity allo- cation for layer-wise pruning of large language models
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 63d92d98-cf18-45a9-a45d-6a71f0bcdb09 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a13e492f-863c-42fb-816e-a0528327c55a · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Navigating Extremes: Dynamic Sparsity in Large Output Spaces
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 840f2a7a-b166-4dbd-8102-4b404d28ec5f · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67235692-f973-4818-a960-1f910571583d · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense TrustLLM: Trustworthiness in Large Language Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa89ebd6-d450-44d1-9013-7aa020b7f05f · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 778cdd0d-f1e7-4867-b52e-1cc7557956df · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Autodan: Generating stealthy jailbreak prompts on aligned large language models
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7cdadd0f-7f26-40e0-83d7-617cc7654d8e · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Automatically auditing large language mod- els via discrete optimization
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 894c6b5b-92f3-4d3d-9ee8-b8e810b3b87f · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Making them ask and an- swer: Jailbreaking large language models in few queries via disguise and reconstruction
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5df229dc-92c6-4790-b0ac-a7dcdadbf21b · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Beavertails: Towards im- proved safety alignment of llm via a human-preference dataset
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 334cc4f1-54fd-45de-949c-8bb81b5e6779 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Pointer Sentinel Mixture Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e3347ca-54f3-4957-8d95-4b6af5c0d62e · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Measuring Massive Multitask Language Understanding
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00071c23-c419-4149-b7ba-be718d7fb249 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Qlora: Efficient finetuning of quan- tized llms
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation eb141789-1fd9-4c8f-895b-2aebe0be15ac · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccb460b1-b160-421d-a89a-934d83896715 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Mixtral of Experts
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bd9bd37-2bfb-4474-83a5-fff8d3c9e48c · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Zephyr: Direct Distillation of LM Alignment
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faabd1bf-74b9-4db4-8168-81affad895e7 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Qwen2.5 Technical Report
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a5aa9ea-1b8c-4af4-8fb0-b5de3faa4aee · outbound
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8efe3443-9481-48b6-a96d-ff673f06b509 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Language models are few-shot learn- ers
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a7600c1-c273-4d04-9415-e4c8c40dab33 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 553ca044-2223-4d81-820c-df4a5472aaea · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8404c6c5-62b5-44a5-b9bd-9a2607544e79 · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Towards Deep Learning Models Resistant to Adversarial Attacks
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f45c01d-dfe1-44ed-96e7-a634fb74bb6f · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Stanford alpaca: An instruction- following llama model, 2023
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fa618968-8898-4dca-a5ae-b76e4f6f171f · outbound
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Unresolved cited work
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 80a66c5d-3841-47f2-9391-cd8c2d36cf8c · inbound
Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c9d6148-5e6c-4d0a-b445-f368849463d7 · inbound
Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.