Pith. sign in

Paper Citation Record · LEDGER

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense

As of 10 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 2 inbound Pith citation observations for arXiv:2502.00840.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00840 v2

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:37:39.675692Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:00:19.630547Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T23:10:21.867778Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e2140333-e41d-4262-babf-344189700f7e · outbound

This paper cites Chat generative pre-trained transformer (chat- gpt).

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Chat generative pre-trained transformer (chat- gpt)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.445405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.445405Z digest=sha256:42ee88a13276d6222f60f9499298da6ce23b41b3013e39440dabc94faf28bdce

Observation 0cea145c-5b3c-45de-a31e-bd36475a504c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.449262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.449262Z digest=sha256:5c32680efb806aa24e4718ee96f4aa0e3d3f2bb345f3083f95fe21c232d55e1f

Observation 5ba827a8-8fac-4239-962e-599a67040617 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Gemma: Open Models Based on Gemini Research and Technology

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.452541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.452541Z digest=sha256:c2e4b4ebfeb14096a6244bc2d7ad441193e97a8ef0bfa759d43e17db694d90ed

Observation 74610e66-12fc-4343-8437-0fcfe916ef52 · outbound

This paper cites Mistral 7B.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Mistral 7B

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.455573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.455573Z digest=sha256:91b6a89887fe333b0fd683f28363f5a32413a1d6501415b246a5088a50d0d20b

Observation 01e4ab5a-08e7-4396-9012-3dfd8245a74f · outbound

This paper cites The Falcon Series of Open Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense The Falcon Series of Open Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.458365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.458365Z digest=sha256:6edff8fd48e5ec4143c9e093ae8d134014e1edb9da0bb9b53d6033d1ccbffee7

Observation 64401fd9-c23c-4d59-aa32-5ca14ea0d8f3 · outbound

This paper cites Qwen Technical Report.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Qwen Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.461305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.461305Z digest=sha256:e0d27f3d92bae9dfea7b1445671f68dce8b8aef8912c2741ebbb75a7ed10edb2

Observation d0ec0fee-50b6-4ebe-92af-222ca99d8f52 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Gaussian Error Linear Units (GELUs)

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.464411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.464411Z digest=sha256:4d29c94dfe7b7eaf3613a63a83046e0de70e0908a11e62b7edc7fb47d6816ee3

Observation c704735d-6f10-4dcb-b2a8-67c736845040 · outbound

This paper cites Sigmoid- weighted linear units for neural network function ap- proximation in reinforcement learning.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sigmoid- weighted linear units for neural network function ap- proximation in reinforcement learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.467467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.467467Z digest=sha256:0119007caf342ea973e04add22d86724e8d350739c63f24bad263bf8a5f073c9

Observation a6aaae0c-d6cd-4a15-bb84-41a7bbc6a0cb · outbound

This paper cites Awq: Activation- aware weight quantization for on-device llm compres- sion and acceleration.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Awq: Activation- aware weight quantization for on-device llm compres- sion and acceleration

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.470224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.470224Z digest=sha256:e783eae58fea9d13862cef9ddfbc1eb570e92641434a6361968e25b4ecc5578b

Observation 35836eb3-c6d7-437f-8d81-84cfea072082 · outbound

This paper cites GPTQ: Accurate post-training compression for generative pretrained transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense GPTQ: Accurate post-training compression for generative pretrained transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.472818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.472818Z digest=sha256:cdd737ce4bd5e165466eca0f6627628f94bb07cf1c4a476dfded39439bba92d7

Observation 479f9a15-8326-4792-8bfd-ec4a0f8371c4 · outbound

This paper cites Billm: Pushing the limit of post-training quan- tization for llms.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Billm: Pushing the limit of post-training quan- tization for llms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.475289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.475289Z digest=sha256:62fd9137d1528492d198a67cdd7a1b3bcd2e71478ae1674773223e59f2c8619b

Observation 651d0dfb-4ab2-45ee-9c83-d4503e1b4c98 · outbound

This paper cites Sparsegpt: Massive lan- guage models can be accurately pruned in one-shot.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sparsegpt: Massive lan- guage models can be accurately pruned in one-shot

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.477774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.477774Z digest=sha256:f3cfe4fc4a1860e7690d6854a7562467e58d940f51fa8255bafe41c162b7e9ff

Observation 2bbb1169-9b05-45db-bc95-7c1ca97e02af · outbound

This paper cites Llm- pruner: On the structural pruning of large language mod- els.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Llm- pruner: On the structural pruning of large language mod- els

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.480322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.480322Z digest=sha256:fb877a6d9ddebf86785dd09c81381db2c75333013ecbab7ebb520ed47a56cd50

Observation d7a430ef-fed4-41b6-b1ba-6de630ba7c3e · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.482768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.482768Z digest=sha256:107c98bc89e8501f7854466601d698b45562c0c4b7739c74f0e29fb59b47d113

Observation faed12e6-880d-4a75-9f5a-8bbbb02903de · outbound

This paper cites Structured Pruning of Large Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Structured Pruning of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.485481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.485481Z digest=sha256:afdf3fe20f001e9f94cbec842ac6589ea23ef968044abe3e9c88e2301866d445

Observation 058c99be-1b39-4bf5-9a2b-a218ac461f99 · outbound

This paper cites Fast inference from transformers via speculative decoding.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Fast inference from transformers via speculative decoding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.488230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.488230Z digest=sha256:4f571fc0d7f69b11718bfa7f07935002fd331f62e92b088628ae3b4d931086d7

Observation 68d3aff6-70f9-42cc-b2b5-c48b44aca84b · outbound

This paper cites Speculative decoding with big little de- coder.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Speculative decoding with big little de- coder

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.490567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.490567Z digest=sha256:0799f15efa387b718bc13be76f3ea178895bb0fa51e97b9ea37f7478ae1cfe8c

Observation bdda5e25-d4fe-4fb3-92b7-8fe140a74b12 · outbound

This paper cites Iron: Private infer- ence on transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Iron: Private infer- ence on transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.492867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.492867Z digest=sha256:2559b3236720a4157504992805ce902ac2ae0bc55a573ccccd1ff5a1504c8562

Observation a5dafbbe-a6ca-40c3-a41c-8fb73b586f97 · outbound

This paper cites Ciphergpt: Secure two- party gpt inference.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Ciphergpt: Secure two- party gpt inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.494884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.494884Z digest=sha256:4c1fa52e014302765b8034b39e536261dfd88512edb214d1e88c0cd9ad79477c

Observation be552ba8-a521-4ba9-9f1b-57b70edcef59 · outbound

This paper cites Bolt: Privacy-preserving, accu- rate and efficient inference for transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Bolt: Privacy-preserving, accu- rate and efficient inference for transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.496815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.496815Z digest=sha256:29ad9f11753648028ad437bda962b69310222bacee17a62f2abf8b707e0b8cbb

Observation 19367cde-fe3b-4f0e-b017-bbd0653347c2 · outbound

This paper cites BumbleBee: Secure Two-party Inference Frame- work for Large Transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense BumbleBee: Secure Two-party Inference Frame- work for Large Transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.208924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.498787Z digest=sha256:04d1f35dfaced57060fe47afc437655d9c0e232ed6c67ce8c702f5b3f603fb0a

Observation 1b0a96f9-7cf6-43b7-8d4a-bccdc494eec1 · outbound

This paper cites Secure transformer infer- ence made non-interactive.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Secure transformer infer- ence made non-interactive

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.194371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.502964Z digest=sha256:041e1a469f42a7a694dd62955e0b464357e7e77d83ebb8acde682fc19a9f2e72

Observation 4e6321c5-dda7-4236-96aa-7c40970326c2 · outbound

This paper cites Training-Free Activation Sparsity in Large Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Training-Free Activation Sparsity in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.505073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.505073Z digest=sha256:0e5469cd27ff4e6f1ac245ee7ee1f184de3469216565e7403f57d21b54fc8053

Observation 287bf1e2-4352-430b-ac76-5ab02d64ee48 · outbound

This paper cites ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.507833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.507833Z digest=sha256:e87d8a9e44f36af002b3c804c60920176fe625ec97997e2bb46dd0b36b2853f3

Observation dcc82b43-ba2b-4f9a-b7d9-6cbb855b197e · outbound

This paper cites Relu strikes back: Exploiting activation sparsity in large lan- guage models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Relu strikes back: Exploiting activation sparsity in large lan- guage models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.187461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.510440Z digest=sha256:94da57292622b852d6cefa410fc4cec9e683d2ffff6c2aa2b7ce3c2390ce6dff

Observation 288148a7-7b86-42e1-af4f-39648d09b0d8 · outbound

This paper cites Smoothquant: Accu- rate and efficient post-training quantization for large language models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Smoothquant: Accu- rate and efficient post-training quantization for large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.180020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.512941Z digest=sha256:2e31454cf02d91f698c30ac9193697ad90d9c07f0733f3c4a5e1b72ef695fc67

Observation a1d2126a-2aa5-4d7a-869f-5e4a87678d1a · outbound

This paper cites Omniquant: Omnidirection- ally calibrated quantization for large language models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Omniquant: Omnidirection- ally calibrated quantization for large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.172593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.515463Z digest=sha256:a8320bc0a70049be61a0f3b82023eba6a727712835e14258289f06b573d4c94a

Observation 8c840334-b818-4007-91e1-3b4854e5963b · outbound

This paper cites Sirnn: A math library for secure rnn inference.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sirnn: A math library for secure rnn inference

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.165223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.517951Z digest=sha256:9fcf31976a7451a2db492e06efbb6fa744000c408d98873092bc6f3ff8df2081

Observation 4f1e32b9-8c20-42b2-84ef-e4f526db4df2 · outbound

This paper cites From chatgpt to threat- gpt: Impact of generative ai in cybersecurity and privacy.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense From chatgpt to threat- gpt: Impact of generative ai in cybersecurity and privacy

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.157872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.520323Z digest=sha256:079c7cd616c6c4ca79e0e9698dab7103368cf6e458444b8e4f964139015cfbd4

Observation 637b992f-b3f1-45cf-ba58-87c2879cdb14 · outbound

This paper cites Exploiting large language models (llms) through decep- tion techniques and persuasion principles.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Exploiting large language models (llms) through decep- tion techniques and persuasion principles

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.150589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.522839Z digest=sha256:dbbca07ef2ef9990eacc83b422ddc9998aa04999917ebb28425727b97faa173f

Observation 4e4e6dbc-4e9c-4b45-afbe-ce3af4e2d85b · outbound

This paper cites PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.525244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.525244Z digest=sha256:f51fe0a032e7e78b0251c9dc3cee9125241d63a3b4754d3b11c5fe243a2b2378

Observation 1ef5ecec-76c2-49c2-9151-df6e15127c28 · outbound

This paper cites Training language models to follow instructions with human feedback.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Training language models to follow instructions with human feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.527855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.527855Z digest=sha256:7bf817acd9178363cb1c87d537fceebd05507e415940a24104cbcfc0a7a915b9

Observation c76fa3c1-23e2-4af4-95c2-8bb3a4510ec2 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.530426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.530426Z digest=sha256:19efd5f853de246f15d8dbc7c1f9cbf132db5c15115b2a945342763363f92298

Observation 990fe8fb-ecd8-4a21-a561-8001f009b155 · outbound

This paper cites Di- rect preference optimization: Your language model is secretly a reward model.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Di- rect preference optimization: Your language model is secretly a reward model

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.139508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.533289Z digest=sha256:47c8a7da1885438eed614a0dfc47f071e4ad39b242bce9f5015c1feb556ee9f5

Observation 28292dea-c199-4d8d-b0c6-068dbfc45920 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.535809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.535809Z digest=sha256:7a0efc27912b15d1ae3abe488271652ddc233d9f65f3d2756dd8a717771b9489

Observation 21dcb9ae-d6d4-444f-96b0-4f402b7abfbe · outbound

This paper cites Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.538531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.538531Z digest=sha256:3f92427f05aa1689514483a7f11b253a38f4a6ddc8a59aa41ccfab1115f71b6a

Observation 648e5c71-b1cd-46a7-b4ce-4f903832b0df · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.541244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.541244Z digest=sha256:26aa870c40b22247cd3fafdf834c8e22dd5fc4d442d6ceeed958db703a4f966f

Observation 554ca386-87ae-4910-a4bb-96a9a97df6b0 · outbound

This paper cites A simple and effective pruning approach for large lan- guage models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense A simple and effective pruning approach for large lan- guage models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.132421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.544127Z digest=sha256:9aae2541c44f67a64a9dca7985e398fd5fd394eead1f2049fc7838b6a448d83f

Observation eb45a763-0ab2-44c5-be77-10d227b55ad4 · outbound

This paper cites Dynamic context pruning for efficient and interpretable autoregressive transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Dynamic context pruning for efficient and interpretable autoregressive transformers

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.125481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.546825Z digest=sha256:df5951ecd157b09de042b7970df69a619e64c6712dcf99eede8bf1879cf439e7

Observation 7ce3b523-7318-4acf-8037-686931756f86 · outbound

This paper cites Adapting language models to compress contexts.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Adapting language models to compress contexts

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.118461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.549206Z digest=sha256:afde7e8027224fecca9d3fa37f6281e92887f571fce6063019a9d12eabaeda34

Observation c805132b-81c1-4aa9-86dc-05ae00074791 · outbound

This paper cites {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache manage- ment.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache manage- ment

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.111328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.551730Z digest=sha256:078d265bd7e04321c0460c5c3870ff50bf75258d7286c595f277ebd948cc79bd

Observation 71451ffd-5c83-4a1e-a185-f5420ad3677c · outbound

This paper cites {IMPRESS}: An {Importance- Informed}{Multi-Tier} prefix {KV} storage system for large language model inference.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense {IMPRESS}: An {Importance- Informed}{Multi-Tier} prefix {KV} storage system for large language model inference

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.104171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.554120Z digest=sha256:e0aab346cbb7f29d1e1bfd3161c84004e89e7dff114427fab14d6bc61946628c

Observation 3a1c35c4-aa16-4794-9217-f105d484939a · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.098118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.556605Z digest=sha256:977913708eafc668523d6311e59d8dbd6e99cf17aafd3300a97ff6076ae28476

Observation e7788879-bce4-4f23-b5e5-cc32f49dd2cb · outbound

This paper cites Kvquant: Towards 10 million con- text length llm inference with kv cache quantization.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Kvquant: Towards 10 million con- text length llm inference with kv cache quantization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.091674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.559046Z digest=sha256:ab78751255f9b236d7df0ff030f9866dfec8786b75d56dc7e5c7a22653de0fe2

Observation 10aad630-7787-44c9-855c-ac005805c07b · outbound

This paper cites Sequoia: Scalable and robust speculative de- coding.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sequoia: Scalable and robust speculative de- coding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.085231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.561483Z digest=sha256:2bfb15e44e3baf97fa68c5823678b062e93ed8f30ef2d44ee0ada44c42becb1a

Observation 48c61911-7bba-476a-bea3-44a973b2df16 · outbound

This paper cites Decoding compressed trust: Scrutinizing the trustworthiness of efficient llms under compression.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Decoding compressed trust: Scrutinizing the trustworthiness of efficient llms under compression

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.078252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.563904Z digest=sha256:337f28f278be2c71a21808c5312dd67245b0d8700ebcb95424061111782246ed

Observation 45a28834-227a-4d5b-8201-4e64193306de · outbound

This paper cites Harmlevelbench: Evaluating harm-level compliance and the impact of quantization on model alignment.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Harmlevelbench: Evaluating harm-level compliance and the impact of quantization on model alignment

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.070834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.566243Z digest=sha256:8251bbd57399983c530a1cb165438b64316d686aff66890e4ed2d75afd3a2c66

Observation 5b8a9931-e89a-415e-892e-742fb4724cb2 · outbound

This paper cites Exploiting llm quantization.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Exploiting llm quantization

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.063577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.568667Z digest=sha256:51ea8a2168f60694522ca67dcc969f68db3e437437b2a7dd56f6b07c843cf838

Observation ceffec12-a271-4d19-b0cb-f420f35f503f · outbound

This paper cites Zeroquant: Efficient and affordable post-training quantization for large-scale transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Zeroquant: Efficient and affordable post-training quantization for large-scale transformers

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.056232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.571296Z digest=sha256:082eed60e7b5773e602c82277f74f5cb64b1e4f8f8576261f7e3860985a57664

Observation 58b43f76-c1a2-4ce1-b0e4-34a321323825 · outbound

This paper cites Qllm: Accurate and efficient low-bitwidth quantization for large language models.In- ternational Conference on Machine Learning, 2024.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Qllm: Accurate and efficient low-bitwidth quantization for large language models.In- ternational Conference on Machine Learning, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.048307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.573726Z digest=sha256:03d1079af2bbdd362b55904047d7589224d066834d048f371bc1b8f7e54dab5c

Observation 4b9e6550-0094-4b81-ad21-6164bfb5beee · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Introducing meta llama 3: The most capable openly available llm to date

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.575830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.575830Z digest=sha256:37f8d468bfc143ace0e20b86bbf652de9c6cd61ee7b423f212f56b295c51424d

Observation 036314c3-df4a-428e-85ef-917363701540 · outbound

This paper cites Mpcformer: fast, performant and private transformer inference with mpc.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Mpcformer: fast, performant and private transformer inference with mpc

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.037095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.577796Z digest=sha256:63a812bac2381c0423a6f4b382b4441e42dc69d0206ea8bfa0b8622e0a2c192c

Observation bdb1cc3d-e089-4a1e-9e20-bf96c91f299b · outbound

This paper cites Encryption-Friendly LLM Architecture.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Encryption-Friendly LLM Architecture

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.579915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.579915Z digest=sha256:caf908e98cfb629d8d88263da05436364965aebcff994bcf4ef807a6faa42739

Observation 9f2ac483-69c7-42ab-bbfb-31c9b162cb94 · outbound

This paper cites Deja vu: Con- textual sparsity for efficient llms at inference time.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Deja vu: Con- textual sparsity for efficient llms at inference time

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.029642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.582055Z digest=sha256:5794c41a684a9eaba1051403f051983e10d44b0c9a36eff715d91ee7dbec3122

Observation 6da7ef52-b474-4671-86df-2cb90b6c996c · outbound

This paper cites Cats: Context-aware thresh- olding for sparsity in large language models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Cats: Context-aware thresh- olding for sparsity in large language models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.022246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.584000Z digest=sha256:ecebba139462dd688081d0e7e1bf22516a13027848339006b92f81c8b1f3c8e9

Observation d9906681-c789-4184-9a25-d9900b8443bc · outbound

This paper cites Judging llm-as- a-judge with mt-bench and chatbot arena.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Judging llm-as- a-judge with mt-bench and chatbot arena

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.586311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.586311Z digest=sha256:2537f3f77e6ec250274fadf45d025ac41a52cf4f46ee53ef70057f483a24ddb9

Observation 04a0dd48-3eb6-413b-94ef-63469149a2a4 · outbound

This paper cites Assessing the brittleness of safety alignment via pruning and low-rank modifica- tions.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Assessing the brittleness of safety alignment via pruning and low-rank modifica- tions

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.010797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.588409Z digest=sha256:d29ec2eeef21f456f3343ed9bacbaf0e4e9f6bc30f95376304cc2059decfba7a

Observation f9eaa58c-f081-48d1-9de3-5221b7b1162b · outbound

This paper cites Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.590442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.590442Z digest=sha256:3de30172a536b2e918f3c3161efa7545a4a53bb286b1ce679646f567303c1e8e

Observation b32f45a2-16e6-49de-8886-851d2c32f982 · outbound

This paper cites Lisa: Lazy safety alignment for large language models against harmful fine-tuning at- tack.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Lisa: Lazy safety alignment for large language models against harmful fine-tuning at- tack

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.003409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.593006Z digest=sha256:092b147b7fe3b9f0892f6638bd4bc7f92f162e8d0ab159fc0c41a4634e77b2b4

Observation a5fb86ba-5cd5-4115-adb7-184980feab14 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.595425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.595425Z digest=sha256:791511ebd73bce52c3e4c0f4e106b80286667d9f1b1bc879de33ce671de4a277

Observation a2012904-6ca2-4b52-8886-23136f88c45a · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.598235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.598235Z digest=sha256:823bc43e1bd648abc6ffa3c3346d2fddcb43d642b810e600855310ccd9b7e2ae

Observation e98b3c49-8494-4c98-9d5a-0b2b8460beac · outbound

This paper cites Detoxifying Large Language Models via Knowledge Editing.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Detoxifying Large Language Models via Knowledge Editing

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.600920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.600920Z digest=sha256:e6b641aa536bfeac3d7a586a729e089051fd4a16cdfdd3488a94a3b9933bc1b6

Observation 9ed5edc6-4186-4910-8e00-71df8537532c · outbound

This paper cites Recovering the Pre-Fine-Tuning Weights of Generative Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Recovering the Pre-Fine-Tuning Weights of Generative Models

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-09T17:37:39.818906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.603681Z digest=sha256:3f34eeb00474cc723eaacd113f04b0b391b40b4b0381e69d316cc04d4baffb92

Observation dcecaae4-378e-4c13-a3a7-3cc5ef393661 · outbound

This paper cites Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.606386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.606386Z digest=sha256:1737abc8755f5631c08877ade514f48b96c74d090129e8f46bdf0069614283ae

Observation 0d6d80ec-20af-4486-bbc5-73332f878ca0 · outbound

This paper cites Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.609496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.609496Z digest=sha256:db948e6c0b0c9d010dc26af0c791312e1645a181f81d02b1b9989b30e1a4098a

Observation c1a338de-b2c4-4f79-8bcc-dacded386d85 · outbound

This paper cites SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.612801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.612801Z digest=sha256:82510511565d361d2c7c4d1750dbc8635df3c58676319d45f295e24dfd220e7d

Observation 6a15da23-7dd0-4138-8b6c-b2475b3ffc78 · outbound

This paper cites Compressing LLMs: The Truth is Rarely Pure and Never Simple.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Compressing LLMs: The Truth is Rarely Pure and Never Simple

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.615850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.615850Z digest=sha256:3d02f4bb624ef79a5e1503e313058fd46298d598b7363ccf466710e8b93d11e8

Observation 960c1af9-ac41-4784-877c-0b3489ae88f3 · outbound

This paper cites Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.618406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.618406Z digest=sha256:a8303909db9a63dfe675a6d13c8a066421953d335785773d59bafb75037d5b2b

Observation d45dd858-0691-4600-b9dd-4eca0d8c4165 · outbound

This paper cites Discovering sparsity allo- cation for layer-wise pruning of large language models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Discovering sparsity allo- cation for layer-wise pruning of large language models

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.996108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.620977Z digest=sha256:1daa637a3c5af63f09ba4321a33a751f1bd2141a2a93f1dd982f85dd9457d011

Observation 63d92d98-cf18-45a9-a45d-6a71f0bcdb09 · outbound

This paper cites Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.623449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.623449Z digest=sha256:e847f506d44ab836b9e7d21b61f47e2af8d0ac80856ac94c2010c158aeb8c659

Observation a13e492f-863c-42fb-816e-a0528327c55a · outbound

This paper cites Navigating Extremes: Dynamic Sparsity in Large Output Spaces.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Navigating Extremes: Dynamic Sparsity in Large Output Spaces

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-09T17:37:39.772417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.626079Z digest=sha256:5501c8d2ecae74346d2a0e41a74e045706f43dff1a8a11ca2b671876c4a08796

Observation 840f2a7a-b166-4dbd-8102-4b404d28ec5f · outbound

This paper cites Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.628959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.628959Z digest=sha256:becb0f3f883f5d40fcfeb67642a8b903f7c2d2e80e9e81134c8a239be1391c1f

Observation 67235692-f973-4818-a960-1f910571583d · outbound

This paper cites TrustLLM: Trustworthiness in Large Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense TrustLLM: Trustworthiness in Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.631778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.631778Z digest=sha256:42d0809fba8dee17155331abb311a5a780aad51c015a0f35b7ac7d0a8240761c

Observation aa89ebd6-d450-44d1-9013-7aa020b7f05f · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.634473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.634473Z digest=sha256:608efac3b316d82218012b014eda68b71db9229b56512ee0451881a416caa0ca

Observation 778cdd0d-f1e7-4867-b52e-1cc7557956df · outbound

This paper cites Autodan: Generating stealthy jailbreak prompts on aligned large language models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Autodan: Generating stealthy jailbreak prompts on aligned large language models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.988585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.637030Z digest=sha256:139bad0ca16ec7eb1b2e58bc2c51538d97c198517f0975997c3020e8c0207106

Observation 7cdadd0f-7f26-40e0-83d7-617cc7654d8e · outbound

This paper cites Automatically auditing large language mod- els via discrete optimization.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Automatically auditing large language mod- els via discrete optimization

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.981980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.639157Z digest=sha256:cef6f6b0a24b29acd9e65e0b0bc8849d0d0c55c6e8bcc2114335b4f1b1502a1f

Observation 894c6b5b-92f3-4d3d-9ee8-b8e810b3b87f · outbound

This paper cites Making them ask and an- swer: Jailbreaking large language models in few queries via disguise and reconstruction.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Making them ask and an- swer: Jailbreaking large language models in few queries via disguise and reconstruction

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.975523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.641111Z digest=sha256:3d0953769992668d36a5de5726372ce0c16516a01712ceb012e0aca7e41ca46a

Observation 5df229dc-92c6-4790-b0ac-a7dcdadbf21b · outbound

This paper cites Beavertails: Towards im- proved safety alignment of llm via a human-preference dataset.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Beavertails: Towards im- proved safety alignment of llm via a human-preference dataset

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.968993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.643119Z digest=sha256:749aff9e2ba9534fcfe0a9a53b5ef0dff75c9c42234ea2ff0cb95e7e0f1d3500

Observation 334cc4f1-54fd-45de-949c-8bb81b5e6779 · outbound

This paper cites Pointer Sentinel Mixture Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Pointer Sentinel Mixture Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.645158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.645158Z digest=sha256:c8a29fe312364d260a45fd803fa015b6a8d72fcb17a128eb8edfdcdba9103d4b

Observation 8e3347ca-54f3-4957-8d95-4b6af5c0d62e · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Measuring Massive Multitask Language Understanding

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.647318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.647318Z digest=sha256:fab3201650159f64db7c7b03e302671c6f00f475b5c7d21475ffedd3ef9f42a3

Observation 00071c23-c419-4149-b7ba-be718d7fb249 · outbound

This paper cites Qlora: Efficient finetuning of quan- tized llms.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Qlora: Efficient finetuning of quan- tized llms

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.961340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.649851Z digest=sha256:a94bc948ad7c920a14b0104aecd00e45d27cb19859939cdef754d784bae93a2a

Observation eb141789-1fd9-4c8f-895b-2aebe0be15ac · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.651866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.651866Z digest=sha256:cf849d6b9ab3022e783ed343595e2070de7b6945720f75cfa6098382cc20a87c

Observation ccb460b1-b160-421d-a89a-934d83896715 · outbound

This paper cites Mixtral of Experts.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Mixtral of Experts

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.654525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.654525Z digest=sha256:4bc9e723a7b4c60acf70f29b8d8e924afb7074174b1e2860bcaa884755e5ccc7

Observation 5bd9bd37-2bfb-4474-83a5-fff8d3c9e48c · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Zephyr: Direct Distillation of LM Alignment

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.657148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.657148Z digest=sha256:f821de813342c7d00112db31b5e2d4bf90c4c0763c3434fa5d99b5232c6e9b44

Observation faabd1bf-74b9-4db4-8168-81affad895e7 · outbound

This paper cites Qwen2.5 Technical Report.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Qwen2.5 Technical Report

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.660075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.660075Z digest=sha256:800f2d3fd946c394167db705062b3146a5201c4a3a475dc18414e402549df1db

Observation 5a5aa9ea-1b8c-4af4-8fb0-b5de3faa4aee · outbound

This paper cites Hashimoto.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Hashimoto

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.662669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.662669Z digest=sha256:6b980a08ee32d4f99f496b7b30776d1cfef74311aeba04ca896446a6df25478a

Observation 8efe3443-9481-48b6-a96d-ff673f06b509 · outbound

This paper cites Language models are few-shot learn- ers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Language models are few-shot learn- ers

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.665104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.665104Z digest=sha256:7ceb8e2ab761de0ce68c978582f7390a87e3fc281b109813f956423b39f0cec4

Observation 4a7600c1-c273-4d04-9415-e4c8c40dab33 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.667643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.667643Z digest=sha256:ddee8b7c0de1f09b9071febddca0fc1f6af64892e38f50d5938b041dfe520a44

Observation 553ca044-2223-4d81-820c-df4a5472aaea · outbound

This paper cites Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.670171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.670171Z digest=sha256:6a6a90b62ad9127f0e560103c4375f38eb0c6f533e1cf08916e76b81bab64a22

Observation 8404c6c5-62b5-44a5-b9bd-9a2607544e79 · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.672901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.672901Z digest=sha256:f1c2f680dc1cfdba15fb46a967a0e78b34bbab301bbc404b470f22d62ef2b683

Observation 5f45c01d-dfe1-44ed-96e7-a634fb74bb6f · outbound

This paper cites Stanford alpaca: An instruction- following llama model, 2023.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Stanford alpaca: An instruction- following llama model, 2023

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.942588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.675692Z digest=sha256:0d8f22e5a7f472fa22691e631be88d8f26d1e10e6106f76d8a451f605301ae35

Observation fa618968-8898-4dca-a5ae-b76e4f6f171f · outbound

This paper cites an unresolved cited work.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:37:40.201564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:37:39.500970Z digest=sha256:e3fd9a7818985d2b30c882ca3182adfb474df394db4257679d03cb0cc13c4683

Pith citing papers

Observation 80a66c5d-3841-47f2-9391-cd8c2d36cf8c · inbound

Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models cites this paper.

Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:19.630547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:19.630547Z digest=sha256:7918730f5e92091fca15042a251074b5f6dda4d2db7f7750c06aa5c01dd3a2ec

Observation 4c9d6148-5e6c-4d0a-b445-f368849463d7 · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-04T23:10:21.874478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T23:10:21.257150Z digest=sha256:8ccb87dd7a713b7c0d0183a0423140854d12264c96dfb38e10559195215f292e