Pith. sign in

Paper Citation Record · LEDGER

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense

As of 10 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 2 inbound Pith citation observations for arXiv:2502.00840.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00840 v2

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:37:39.675692Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:00:19.630547Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T23:10:21.867778Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e2140333-e41d-4262-babf-344189700f7e · outbound

This paper cites Chat generative pre-trained transformer (chat- gpt).

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Chat generative pre-trained transformer (chat- gpt)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.445405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.445405Z digest=sha256:e5243d98f7fa59e1c847245deb310fd00ce633fed83dbd3fe68ef9dd39c9b01b

Observation 0cea145c-5b3c-45de-a31e-bd36475a504c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.449262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.449262Z digest=sha256:881e7fdd614202b6a88ad95e35b697a4aaa42247f77a3eb207c936f50449fd1a

Observation 5ba827a8-8fac-4239-962e-599a67040617 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Gemma: Open Models Based on Gemini Research and Technology

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.452541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.452541Z digest=sha256:7451af118783cdba72cd9cd056a391d19e4e8c3adb726f867dacf84dcd50f29b

Observation 74610e66-12fc-4343-8437-0fcfe916ef52 · outbound

This paper cites Mistral 7B.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Mistral 7B

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.455573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.455573Z digest=sha256:7f178b368d51adb0c426afb207ed5443594c4ec74d1d4bb9bfdee1dd954b253d

Observation 01e4ab5a-08e7-4396-9012-3dfd8245a74f · outbound

This paper cites The Falcon Series of Open Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense The Falcon Series of Open Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.458365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.458365Z digest=sha256:528d8585d0617eca012abeff88dc0f2a63da3dec12bc5315df6dc1a4b672e0d0

Observation 64401fd9-c23c-4d59-aa32-5ca14ea0d8f3 · outbound

This paper cites Qwen Technical Report.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Qwen Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.461305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.461305Z digest=sha256:0a832d5d6e0ef8a4f4b9edc09d5f61ca81cea4c11a2528e3ffd98ce957849e42

Observation d0ec0fee-50b6-4ebe-92af-222ca99d8f52 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Gaussian Error Linear Units (GELUs)

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.464411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.464411Z digest=sha256:b2beed4d19c07a484c675318ea249728c280d409dc5b6116d5d6e285cc94e73b

Observation c704735d-6f10-4dcb-b2a8-67c736845040 · outbound

This paper cites Sigmoid- weighted linear units for neural network function ap- proximation in reinforcement learning.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sigmoid- weighted linear units for neural network function ap- proximation in reinforcement learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.467467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.467467Z digest=sha256:4340361b56a69ae38d1cc460b043613bbbad610f39ba062c19eb8fcb790389d6

Observation a6aaae0c-d6cd-4a15-bb84-41a7bbc6a0cb · outbound

This paper cites Awq: Activation- aware weight quantization for on-device llm compres- sion and acceleration.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Awq: Activation- aware weight quantization for on-device llm compres- sion and acceleration

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.470224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.470224Z digest=sha256:3b002a74452c27231a402386114c7aad78ded4d2b8620329e4b5c648a1902b53

Observation 35836eb3-c6d7-437f-8d81-84cfea072082 · outbound

This paper cites GPTQ: Accurate post-training compression for generative pretrained transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense GPTQ: Accurate post-training compression for generative pretrained transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.472818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.472818Z digest=sha256:04ca2a59a2babf1f58affb51d9623aadc0d4fab2799b40933991f1b9b65b3bab

Observation 479f9a15-8326-4792-8bfd-ec4a0f8371c4 · outbound

This paper cites Billm: Pushing the limit of post-training quan- tization for llms.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Billm: Pushing the limit of post-training quan- tization for llms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.475289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.475289Z digest=sha256:62625d079c754ffd2014d22178254cc4e83e9f30d47fd59ef25a1e8bdd39a63e

Observation 651d0dfb-4ab2-45ee-9c83-d4503e1b4c98 · outbound

This paper cites Sparsegpt: Massive lan- guage models can be accurately pruned in one-shot.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sparsegpt: Massive lan- guage models can be accurately pruned in one-shot

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.477774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.477774Z digest=sha256:b9575e985c4610d0ef3f099576c1f660258af5e77be1d6f280b26dceeb05296f

Observation 2bbb1169-9b05-45db-bc95-7c1ca97e02af · outbound

This paper cites Llm- pruner: On the structural pruning of large language mod- els.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Llm- pruner: On the structural pruning of large language mod- els

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.480322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.480322Z digest=sha256:86c016ce2ae8c5f2973db7e3d2a8d4fffc2aa5c2e6605e0692b2d5d696460bb6

Observation d7a430ef-fed4-41b6-b1ba-6de630ba7c3e · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.482768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.482768Z digest=sha256:83b52f571e5a55644c7e88a3e08b685d0437b58eb0cfe41cb959407dc5b4e57a

Observation faed12e6-880d-4a75-9f5a-8bbbb02903de · outbound

This paper cites Structured Pruning of Large Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Structured Pruning of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.485481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.485481Z digest=sha256:fd4f42b3c06bbdcba37a1a278c41b014730ce3d8bbb2f4cc02ab78fe5a63ef85

Observation 058c99be-1b39-4bf5-9a2b-a218ac461f99 · outbound

This paper cites Fast inference from transformers via speculative decoding.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Fast inference from transformers via speculative decoding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.488230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.488230Z digest=sha256:a473a6947aa0a2f8806f3c440a991266d22aa244bc2f73be1209d237077a73d3

Observation 68d3aff6-70f9-42cc-b2b5-c48b44aca84b · outbound

This paper cites Speculative decoding with big little de- coder.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Speculative decoding with big little de- coder

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.490567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.490567Z digest=sha256:164a49e1cadf1dca3d98bd2e4da8649544a876e02eee7a680b1d8c07f87ce31e

Observation bdda5e25-d4fe-4fb3-92b7-8fe140a74b12 · outbound

This paper cites Iron: Private infer- ence on transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Iron: Private infer- ence on transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.492867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.492867Z digest=sha256:772f23bef53d4be8ef2617c210c3d9b1a7923d66ca9ff0e57c4d7a4829a46809

Observation a5dafbbe-a6ca-40c3-a41c-8fb73b586f97 · outbound

This paper cites Ciphergpt: Secure two- party gpt inference.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Ciphergpt: Secure two- party gpt inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.494884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.494884Z digest=sha256:63dde64dd9e8f699e9f110d80610f7ca2085bbd1cef6a6aef850a8e57ea9c9a4

Observation be552ba8-a521-4ba9-9f1b-57b70edcef59 · outbound

This paper cites Bolt: Privacy-preserving, accu- rate and efficient inference for transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Bolt: Privacy-preserving, accu- rate and efficient inference for transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.496815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.496815Z digest=sha256:c5d5be78d77478d4356b2cbf7ca100b647b810c9cea98b6abdee0e11466a69ee

Observation 19367cde-fe3b-4f0e-b017-bbd0653347c2 · outbound

This paper cites BumbleBee: Secure Two-party Inference Frame- work for Large Transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense BumbleBee: Secure Two-party Inference Frame- work for Large Transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.208924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.498787Z digest=sha256:6ff94a32c47fd4f89d1bf4849d4de2a3b35863b0b0f4c6d255e075692403ea8f

Observation 1b0a96f9-7cf6-43b7-8d4a-bccdc494eec1 · outbound

This paper cites Secure transformer infer- ence made non-interactive.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Secure transformer infer- ence made non-interactive

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.194371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.502964Z digest=sha256:fd22f675d626690aa74563a871b7c0f58725c2e24319f0e300fd3f7595c2d4eb

Observation 4e6321c5-dda7-4236-96aa-7c40970326c2 · outbound

This paper cites Training-Free Activation Sparsity in Large Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Training-Free Activation Sparsity in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.505073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.505073Z digest=sha256:5d6726fe0e321af8097e513637652ea178572d64bac0cf466eec6c0efdc45bd2

Observation 287bf1e2-4352-430b-ac76-5ab02d64ee48 · outbound

This paper cites ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.507833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.507833Z digest=sha256:177d61be547cd8f4b8be23b207d1806aebd387a087b4c4c07384e12517e3913e

Observation dcc82b43-ba2b-4f9a-b7d9-6cbb855b197e · outbound

This paper cites Relu strikes back: Exploiting activation sparsity in large lan- guage models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Relu strikes back: Exploiting activation sparsity in large lan- guage models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.187461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.510440Z digest=sha256:b891a9e97ac78780e7df8060557afe531cc5c20cab43509d75817c5e1d101309

Observation 288148a7-7b86-42e1-af4f-39648d09b0d8 · outbound

This paper cites Smoothquant: Accu- rate and efficient post-training quantization for large language models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Smoothquant: Accu- rate and efficient post-training quantization for large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.180020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.512941Z digest=sha256:271691ee4247a6b2ca2bf43ec97bb938bc7506c5a40e8bd66910d774d23b30a8

Observation a1d2126a-2aa5-4d7a-869f-5e4a87678d1a · outbound

This paper cites Omniquant: Omnidirection- ally calibrated quantization for large language models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Omniquant: Omnidirection- ally calibrated quantization for large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.172593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.515463Z digest=sha256:85cbfe77ce7c5f5fb89ef06fbe4545b52475fe8a0f11351c08b5805b50abc84e

Observation 8c840334-b818-4007-91e1-3b4854e5963b · outbound

This paper cites Sirnn: A math library for secure rnn inference.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sirnn: A math library for secure rnn inference

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.165223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.517951Z digest=sha256:2ec82169e4a009b2f5cb9acb23815d01db3ba269845192c07aefc12f9c6e30c9

Observation 4f1e32b9-8c20-42b2-84ef-e4f526db4df2 · outbound

This paper cites From chatgpt to threat- gpt: Impact of generative ai in cybersecurity and privacy.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense From chatgpt to threat- gpt: Impact of generative ai in cybersecurity and privacy

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.157872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.520323Z digest=sha256:4307e7d2a76c7ae3cb38bfe0d608a107bf21337017b65e736d31ff5525a25519

Observation 637b992f-b3f1-45cf-ba58-87c2879cdb14 · outbound

This paper cites Exploiting large language models (llms) through decep- tion techniques and persuasion principles.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Exploiting large language models (llms) through decep- tion techniques and persuasion principles

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.150589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.522839Z digest=sha256:3ab3294733b38daa697058065a87ad925b62d755df7728090a9e487759b158bc

Observation 4e4e6dbc-4e9c-4b45-afbe-ce3af4e2d85b · outbound

This paper cites PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.525244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.525244Z digest=sha256:ff5dc6853934fc87c30c4eba8452672e81f105139de33eee09186a72e623190c

Observation 1ef5ecec-76c2-49c2-9151-df6e15127c28 · outbound

This paper cites Training language models to follow instructions with human feedback.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Training language models to follow instructions with human feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.527855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.527855Z digest=sha256:4590d707dc796bc9cd138a4e4d59b1fd4854346ee3266822742dafe69e21e82c

Observation c76fa3c1-23e2-4af4-95c2-8bb3a4510ec2 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.530426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.530426Z digest=sha256:d83972f4121b66410bd36bb23d0681e700058693cc9fda76e5f62f83e6b4b437

Observation 990fe8fb-ecd8-4a21-a561-8001f009b155 · outbound

This paper cites Di- rect preference optimization: Your language model is secretly a reward model.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Di- rect preference optimization: Your language model is secretly a reward model

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.139508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.533289Z digest=sha256:a77be52411026aac0dbe2ab8feda0ff6055838c4de2c11d56cd21b04971aa893

Observation 28292dea-c199-4d8d-b0c6-068dbfc45920 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.535809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.535809Z digest=sha256:502e165f5a24aeac149bfc98b4fdd1862486dc680e805d13778186448a041a0a

Observation 21dcb9ae-d6d4-444f-96b0-4f402b7abfbe · outbound

This paper cites Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.538531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.538531Z digest=sha256:c8d4c7481b2bd9ccde054cd8065f1bf973a2619a294ee80f5130bd7cefb45a07

Observation 648e5c71-b1cd-46a7-b4ce-4f903832b0df · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.541244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.541244Z digest=sha256:bf42d3f488dacf6243f57114a95d7f945984757dc74dee50501252992dd2c84f

Observation 554ca386-87ae-4910-a4bb-96a9a97df6b0 · outbound

This paper cites A simple and effective pruning approach for large lan- guage models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense A simple and effective pruning approach for large lan- guage models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.132421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.544127Z digest=sha256:758aee06c3216e72e1e6d093e7a7e3dafa833122485cb926fe412e53610e7b13

Observation eb45a763-0ab2-44c5-be77-10d227b55ad4 · outbound

This paper cites Dynamic context pruning for efficient and interpretable autoregressive transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Dynamic context pruning for efficient and interpretable autoregressive transformers

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.125481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.546825Z digest=sha256:e81210524370febc7e5e68376d04e97cdca1c166084a840dba9d18ebddcc5fd9

Observation 7ce3b523-7318-4acf-8037-686931756f86 · outbound

This paper cites Adapting language models to compress contexts.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Adapting language models to compress contexts

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.118461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.549206Z digest=sha256:0caf86a056f9a894920ddfa4123f6f962c89e387a6c4f64c4b0d03a5a8ae316d

Observation c805132b-81c1-4aa9-86dc-05ae00074791 · outbound

This paper cites {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache manage- ment.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache manage- ment

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.111328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.551730Z digest=sha256:4bd5535bc2edc92ad79ce48524919aa167d9c5a4783caed56e11f98389a3102a

Observation 71451ffd-5c83-4a1e-a185-f5420ad3677c · outbound

This paper cites {IMPRESS}: An {Importance- Informed}{Multi-Tier} prefix {KV} storage system for large language model inference.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense {IMPRESS}: An {Importance- Informed}{Multi-Tier} prefix {KV} storage system for large language model inference

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.104171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.554120Z digest=sha256:6cac8e4ee1cc8bce57e91f060dd520f4786218d58fb76cbcca966c43b7c8dec9

Observation 3a1c35c4-aa16-4794-9217-f105d484939a · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.098118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.556605Z digest=sha256:0a532a1af0c94438452d53d5f7ee0f0f20db9e0a0f4bf3af7bc15658956f2a6f

Observation e7788879-bce4-4f23-b5e5-cc32f49dd2cb · outbound

This paper cites Kvquant: Towards 10 million con- text length llm inference with kv cache quantization.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Kvquant: Towards 10 million con- text length llm inference with kv cache quantization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.091674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.559046Z digest=sha256:2d4bad1901dc5fe3fddc2001c844b6c2ba274678556fc0f94ca8a693b30a8995

Observation 10aad630-7787-44c9-855c-ac005805c07b · outbound

This paper cites Sequoia: Scalable and robust speculative de- coding.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sequoia: Scalable and robust speculative de- coding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.085231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.561483Z digest=sha256:0aa358b168836458396d2c7dba5d01a6298ff17a3e12c1e623783da228797756

Observation 48c61911-7bba-476a-bea3-44a973b2df16 · outbound

This paper cites Decoding compressed trust: Scrutinizing the trustworthiness of efficient llms under compression.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Decoding compressed trust: Scrutinizing the trustworthiness of efficient llms under compression

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.078252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.563904Z digest=sha256:4b937e5feb8f3f887971fac6008ad9bd4f5e2b6bd0cdf5718646d47f6e9bf8b3

Observation 45a28834-227a-4d5b-8201-4e64193306de · outbound

This paper cites Harmlevelbench: Evaluating harm-level compliance and the impact of quantization on model alignment.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Harmlevelbench: Evaluating harm-level compliance and the impact of quantization on model alignment

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.070834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.566243Z digest=sha256:f2a20bfaa0dab9e42992595dd34bc29498dec1fc47094f24a07e999d2779dd2e

Observation 5b8a9931-e89a-415e-892e-742fb4724cb2 · outbound

This paper cites Exploiting llm quantization.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Exploiting llm quantization

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.063577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.568667Z digest=sha256:7af9cf1ecff0d2997d9fb04e2aafe4bc14030af99807b9309a1385c8dc04c51c

Observation ceffec12-a271-4d19-b0cb-f420f35f503f · outbound

This paper cites Zeroquant: Efficient and affordable post-training quantization for large-scale transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Zeroquant: Efficient and affordable post-training quantization for large-scale transformers

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.056232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.571296Z digest=sha256:657eacec4857e90e2f9834c5b3f7a4a61d3f764849a8a00629b9c21938e28751

Observation 58b43f76-c1a2-4ce1-b0e4-34a321323825 · outbound

This paper cites Qllm: Accurate and efficient low-bitwidth quantization for large language models.In- ternational Conference on Machine Learning, 2024.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Qllm: Accurate and efficient low-bitwidth quantization for large language models.In- ternational Conference on Machine Learning, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.048307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.573726Z digest=sha256:236cef48208e3970a757d5e3945d971b93782cded5223676a44855ecd7f3fafe

Observation 4b9e6550-0094-4b81-ad21-6164bfb5beee · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Introducing meta llama 3: The most capable openly available llm to date

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.575830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.575830Z digest=sha256:599188d9e5b658ca7977a9ce2324b222f1f6d170b868c391cc1862a7e49db3d2

Observation 036314c3-df4a-428e-85ef-917363701540 · outbound

This paper cites Mpcformer: fast, performant and private transformer inference with mpc.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Mpcformer: fast, performant and private transformer inference with mpc

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.037095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.577796Z digest=sha256:2055b9fd0eba98939d5d482ac93d34b700d2df658186a8ba0e79dd9a6bd56c70

Observation bdb1cc3d-e089-4a1e-9e20-bf96c91f299b · outbound

This paper cites Encryption-Friendly LLM Architecture.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Encryption-Friendly LLM Architecture

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.579915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.579915Z digest=sha256:721dcccf342a238d07ed657ad9d68a48fbacd3bcc6b61be1210a28724feed46f

Observation 9f2ac483-69c7-42ab-bbfb-31c9b162cb94 · outbound

This paper cites Deja vu: Con- textual sparsity for efficient llms at inference time.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Deja vu: Con- textual sparsity for efficient llms at inference time

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.029642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.582055Z digest=sha256:09c816e3a276b5edde89e2530645be3faa391275630230ed45140f4623babcdb

Observation 6da7ef52-b474-4671-86df-2cb90b6c996c · outbound

This paper cites Cats: Context-aware thresh- olding for sparsity in large language models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Cats: Context-aware thresh- olding for sparsity in large language models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.022246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.584000Z digest=sha256:8a25f891d31a9cefda2b0d9dae1948e073d5143b0ea914f3ac1e76b1728ed9c3

Observation d9906681-c789-4184-9a25-d9900b8443bc · outbound

This paper cites Judging llm-as- a-judge with mt-bench and chatbot arena.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Judging llm-as- a-judge with mt-bench and chatbot arena

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.586311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.586311Z digest=sha256:045f5158061cf85466a34e34d9edc32384c8a5a957c5138ab709c90fed9cec3a

Observation 04a0dd48-3eb6-413b-94ef-63469149a2a4 · outbound

This paper cites Assessing the brittleness of safety alignment via pruning and low-rank modifica- tions.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Assessing the brittleness of safety alignment via pruning and low-rank modifica- tions

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.010797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.588409Z digest=sha256:d75434c89afdbee69a49da2688c6b5490c924e5621a44898434e392b6f77006d

Observation f9eaa58c-f081-48d1-9de3-5221b7b1162b · outbound

This paper cites Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.590442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.590442Z digest=sha256:6f0ec5894529f0f925edd40961799d4511a76edbc6658908c812ed8342eb1476

Observation b32f45a2-16e6-49de-8886-851d2c32f982 · outbound

This paper cites Lisa: Lazy safety alignment for large language models against harmful fine-tuning at- tack.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Lisa: Lazy safety alignment for large language models against harmful fine-tuning at- tack

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.003409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.593006Z digest=sha256:ed78417e43e7acc79d02481423c27a54a868fb338b5f006c94df4ae977fb93ce

Observation a5fb86ba-5cd5-4115-adb7-184980feab14 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.595425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.595425Z digest=sha256:44029ba95a2038d5cd111c4794b324dd642f3c67097782897e1792f7f762d5d8

Observation a2012904-6ca2-4b52-8886-23136f88c45a · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.598235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.598235Z digest=sha256:b82d3ffa8a019e386e6b26cecee47c18ddfbc01fcafcabaee9f6e6c3f646ec8d

Observation e98b3c49-8494-4c98-9d5a-0b2b8460beac · outbound

This paper cites Detoxifying Large Language Models via Knowledge Editing.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Detoxifying Large Language Models via Knowledge Editing

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.600920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.600920Z digest=sha256:422788961fa4b121eedf33057ae9e1eeca6fd4d46669fca6fd11644c078717db

Observation 9ed5edc6-4186-4910-8e00-71df8537532c · outbound

This paper cites Recovering the Pre-Fine-Tuning Weights of Generative Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Recovering the Pre-Fine-Tuning Weights of Generative Models

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-09T17:37:39.818906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.603681Z digest=sha256:7e74ba7e605be74b849776b582b423eda0bef8928c8fd439b46560c1b2e2c5e0

Observation dcecaae4-378e-4c13-a3a7-3cc5ef393661 · outbound

This paper cites Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.606386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.606386Z digest=sha256:a6d48b6105895287df4c4c5d9ba40f6e4bac8ff352f9fa8bfb329ee6b66c61cd

Observation 0d6d80ec-20af-4486-bbc5-73332f878ca0 · outbound

This paper cites Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.609496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.609496Z digest=sha256:6b8b48f27a9c98985423347580875fafc4c0ad1cdf778f7bf9192e834b3cf7bb

Observation c1a338de-b2c4-4f79-8bcc-dacded386d85 · outbound

This paper cites SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.612801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.612801Z digest=sha256:1744c973f7ee0a7926f3a605f4f096af05fa6d30fc6266df2cf3727b93ffd80d

Observation 6a15da23-7dd0-4138-8b6c-b2475b3ffc78 · outbound

This paper cites Compressing LLMs: The Truth is Rarely Pure and Never Simple.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Compressing LLMs: The Truth is Rarely Pure and Never Simple

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.615850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.615850Z digest=sha256:4ea6a9f53e882a9bd3f64ccb82a71df8e5026c02b99b08d93365167d04040d62

Observation 960c1af9-ac41-4784-877c-0b3489ae88f3 · outbound

This paper cites Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.618406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.618406Z digest=sha256:46b20ce56f38dfbfac4316f668915428057ae6ab1da401563a4b097fe6b25f89

Observation d45dd858-0691-4600-b9dd-4eca0d8c4165 · outbound

This paper cites Discovering sparsity allo- cation for layer-wise pruning of large language models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Discovering sparsity allo- cation for layer-wise pruning of large language models

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.996108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.620977Z digest=sha256:ccc0a05def31f0697f9a18acd237fdc08199529a99083fa4593845055de24025

Observation 63d92d98-cf18-45a9-a45d-6a71f0bcdb09 · outbound

This paper cites Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.623449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.623449Z digest=sha256:6d84b09ebefdb57f3e799b7b8aed31b36a9f30afe7842332fa1bab51c64255ee

Observation a13e492f-863c-42fb-816e-a0528327c55a · outbound

This paper cites Navigating Extremes: Dynamic Sparsity in Large Output Spaces.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Navigating Extremes: Dynamic Sparsity in Large Output Spaces

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-09T17:37:39.772417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.626079Z digest=sha256:95af182f174a21a6352b8c684c0f23ffa46eccd2f536cff516a4fc8f7d61b2eb

Observation 840f2a7a-b166-4dbd-8102-4b404d28ec5f · outbound

This paper cites Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.628959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.628959Z digest=sha256:f2b8270d03c86b01048756127e737e252c88b675e2f0460c3124d0c924d770ad

Observation 67235692-f973-4818-a960-1f910571583d · outbound

This paper cites TrustLLM: Trustworthiness in Large Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense TrustLLM: Trustworthiness in Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.631778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.631778Z digest=sha256:be79c6d05bb9259f14a85f737258b402a0dbd22f1ef3d4784feddc55653bc568

Observation aa89ebd6-d450-44d1-9013-7aa020b7f05f · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.634473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.634473Z digest=sha256:2bdc59878dbba8c31196895002180a531e57b9a8661aff16b051b6e5f50e6d55

Observation 778cdd0d-f1e7-4867-b52e-1cc7557956df · outbound

This paper cites Autodan: Generating stealthy jailbreak prompts on aligned large language models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Autodan: Generating stealthy jailbreak prompts on aligned large language models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.988585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.637030Z digest=sha256:2d63ce1763545a266708d7b85caeda72c187a27253bc4ba1f5126286473950f2

Observation 7cdadd0f-7f26-40e0-83d7-617cc7654d8e · outbound

This paper cites Automatically auditing large language mod- els via discrete optimization.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Automatically auditing large language mod- els via discrete optimization

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.981980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.639157Z digest=sha256:decabed9b8b18ebfe42deb2f6e9d091b51c5dac006c57acc0cd1aa765cec8965

Observation 894c6b5b-92f3-4d3d-9ee8-b8e810b3b87f · outbound

This paper cites Making them ask and an- swer: Jailbreaking large language models in few queries via disguise and reconstruction.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Making them ask and an- swer: Jailbreaking large language models in few queries via disguise and reconstruction

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.975523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.641111Z digest=sha256:4b3ed17bc29d517a7e3f660a0c0fa488ec14ae2aa17d5483eb76a5b625383fcd

Observation 5df229dc-92c6-4790-b0ac-a7dcdadbf21b · outbound

This paper cites Beavertails: Towards im- proved safety alignment of llm via a human-preference dataset.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Beavertails: Towards im- proved safety alignment of llm via a human-preference dataset

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.968993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.643119Z digest=sha256:d8f4ccef87d74f9d13c9c75ef5291f7c6f7744f9881acc0f360af079ed480d55

Observation 334cc4f1-54fd-45de-949c-8bb81b5e6779 · outbound

This paper cites Pointer Sentinel Mixture Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Pointer Sentinel Mixture Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.645158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.645158Z digest=sha256:3beb1da205d4a036e99e4bbb7a72211ce823e9d3ece279c1ca3e12792739fde6

Observation 8e3347ca-54f3-4957-8d95-4b6af5c0d62e · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Measuring Massive Multitask Language Understanding

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.647318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.647318Z digest=sha256:afa9032029726a349e6c139cc6781aef41bdc8143d162aa9db083eeb4008d40b

Observation 00071c23-c419-4149-b7ba-be718d7fb249 · outbound

This paper cites Qlora: Efficient finetuning of quan- tized llms.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Qlora: Efficient finetuning of quan- tized llms

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.961340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.649851Z digest=sha256:e8d00b04fb98723ce2e1b7eb93e787d1137a7ff41efeb8afdd0fa307374534ad

Observation eb141789-1fd9-4c8f-895b-2aebe0be15ac · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.651866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.651866Z digest=sha256:3c788789f1eb678320e91bedec434b1e7b2739c9d441bac1f1a9876b1a8277a9

Observation ccb460b1-b160-421d-a89a-934d83896715 · outbound

This paper cites Mixtral of Experts.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Mixtral of Experts

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.654525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.654525Z digest=sha256:d5c8f189f417ef720a7c416e3b4f1c9ef779cb5e9b54b8ad386cf40d7f19eea9

Observation 5bd9bd37-2bfb-4474-83a5-fff8d3c9e48c · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Zephyr: Direct Distillation of LM Alignment

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.657148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.657148Z digest=sha256:7a5e8261a025dff099a034da6eec4f1285273b4c115b30caadbeb69f3f738bf2

Observation faabd1bf-74b9-4db4-8168-81affad895e7 · outbound

This paper cites Qwen2.5 Technical Report.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Qwen2.5 Technical Report

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.660075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.660075Z digest=sha256:4f838cb42cb1f7c4220b24767eac1b37f968cf4d17825276c12e7924d9256744

Observation 5a5aa9ea-1b8c-4af4-8fb0-b5de3faa4aee · outbound

This paper cites Hashimoto.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Hashimoto

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.662669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.662669Z digest=sha256:d547d2a6ff0d611d9b62925b6883826c67c9977e4b9ea3e2ed8dbaeafe3ad93b

Observation 8efe3443-9481-48b6-a96d-ff673f06b509 · outbound

This paper cites Language models are few-shot learn- ers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Language models are few-shot learn- ers

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.665104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.665104Z digest=sha256:fdbc294e0a920831a58e4a28c61d7616872c3e991dc6dbd5dc90e57bd69b71e0

Observation 4a7600c1-c273-4d04-9415-e4c8c40dab33 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.667643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.667643Z digest=sha256:fd650d95a02a77fd7c8b0cb331b7685cda7771d6a6ea1619faf662d2cd4bf30d

Observation 553ca044-2223-4d81-820c-df4a5472aaea · outbound

This paper cites Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.670171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.670171Z digest=sha256:4bb36dddd35fcbb5edabd6a53328825b6b86f31c743e3986d0ebd5a38d2740da

Observation 8404c6c5-62b5-44a5-b9bd-9a2607544e79 · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.672901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.672901Z digest=sha256:dca4b5b3daf01233fb78787e5c80acc14252efdcaa7d5e794bc3f356ec3ec61d

Observation 5f45c01d-dfe1-44ed-96e7-a634fb74bb6f · outbound

This paper cites Stanford alpaca: An instruction- following llama model, 2023.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Stanford alpaca: An instruction- following llama model, 2023

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.942588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.675692Z digest=sha256:d54f138901ea6190883c6873b79791d200821a66e4197dfcfdf539c29112ca50

Observation fa618968-8898-4dca-a5ae-b76e4f6f171f · outbound

This paper cites an unresolved cited work.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:37:40.201564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T17:37:39.500970Z digest=sha256:7307cd97f36052ff420aff155929f402db97e9a681e47c85e6dfd8b652de3ed6

Pith citing papers

Observation 80a66c5d-3841-47f2-9391-cd8c2d36cf8c · inbound

Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models cites this paper.

Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:19.630547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:19.630547Z digest=sha256:c9936733e51cc997a2a89926d0228f0431b5b6fceb88f4e54edc02f444e4b807

Observation 4c9d6148-5e6c-4d0a-b445-f368849463d7 · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-04T23:10:21.874478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:10:21.257150Z digest=sha256:52a4f2a8ead7f0ec1e7c67cb19425b3856064479a22aae0b6a33d85c49144d6f