Pith. sign in

Paper Citation Record · LEDGER

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks

As of 12 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2608.01414.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01414 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:18:53.661855Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1f600858-8d64-49c6-a2d6-d67d271eb025 · outbound

This paper cites Phi-4 Technical Report.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Phi-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:47.641655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:47.641655Z digest=sha256:4161dddb318718ec369f80cbb7113dc0016ff23f15522d12a11e7a060812bd0d

Observation 90e3ad4a-aa29-4c3c-8322-fc053dbcbf88 · outbound

This paper cites Refusal in language models is mediated by a single direction.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Refusal in language models is mediated by a single direction

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:19:01.147174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:47.790026Z digest=sha256:0efd48249e79341c7bd068a4d173f7585a5f2a14b16230e48c6c195b7bcb0f5b

Observation 84f80de4-8a14-447b-9b6a-96163838e204 · outbound

This paper cites Qwen technical report, 2023.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Qwen technical report, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:47.909570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:47.909570Z digest=sha256:efd6139c26f9b33291dc040dbc2526602528565e6e7fc99e47b8c85a21bddc7f

Observation 4074e66f-a830-4527-aeae-d624851a7bd2 · outbound

This paper cites Qwen2.5-vl technical report, 2025.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Qwen2.5-vl technical report, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:48.015481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:48.015481Z digest=sha256:a320c1e8ac9f4b5b61a32422d61753ddd5c6ad05cf1beeb484554c355750474d

Observation 1ce0f009-250b-4a6d-9397-5e55f263ec5f · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:48.167155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:48.167155Z digest=sha256:55be1b8225dac40309722fb2602d374e4678e9cd59ae114de21989ab2b78c6b1

Observation 3c523e77-2bcf-4195-86f6-0f3d87f8611d · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Mechanistic Interpretability for AI Safety -- A Review

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:48.335534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:48.335534Z digest=sha256:16685890503ff49731c4997211e0d8d5b6cfd2f70dd18c89f749738c3700eeee

Observation 3c2733ef-2b62-43f7-8dba-730a14994635 · outbound

This paper cites Language models are Homer simpson! safety re-alignment of fine-tuned language models through task arithmetic.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Language models are Homer simpson! safety re-alignment of fine-tuned language models through task arithmetic

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:48.427125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:48.427125Z digest=sha256:a397cfe4af95378814c623b218c3cca0dcbc7d646c52fba94b597653a3ca3c95

Observation a8e4ed37-5070-4455-b867-09c47f93ea5b · outbound

This paper cites Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:48.516596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:48.516596Z digest=sha256:8ed832156d9d8f6233c0bd9cf0fc65ad6eb5bce96a75b2dca9200a86228a856b

Observation 392568be-bb36-4c6b-a44b-9b8c8ead0abb · outbound

This paper cites Brown et al.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Brown et al

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:19:00.861651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:48.616841Z digest=sha256:2faf595da0cc8619db76acec806685f550f26d6dd94bed23cdb4ffa9503a1c78

Observation a82b5044-2036-4a88-b5c9-3d237b593f6e · outbound

This paper cites Towards understanding safety alignment: A mechanistic perspective from safety neurons.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Towards understanding safety alignment: A mechanistic perspective from safety neurons

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:19:00.512350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:48.747271Z digest=sha256:d5be960ed31aea0f096244c1ac4e6aa6419116bd9845029b3dfeb40ee927d88d

Observation 099a1440-6997-43b3-840a-86ce3e9ea498 · outbound

This paper cites Think you have solved question answering? try ARC, the AI2 reasoning challenge, 2018.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Think you have solved question answering? try ARC, the AI2 reasoning challenge, 2018

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:19:00.155206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:48.913473Z digest=sha256:e7b0b2f080eae2026eee5e376aab56e77003269a09eadd8d7fea303a71a65959

Observation d799eba4-ccc5-4f82-8abb-6a5162de9a0f · outbound

This paper cites Training verifiers to solve math word problems, 2021.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Training verifiers to solve math word problems, 2021

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:49.043158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:49.043158Z digest=sha256:912531ce4e2a2c88389c35335b125b87045b81be2b147559b7e4a912489276a0

Observation 9497005e-7db5-4842-bdde-59636bd6558f · outbound

This paper cites Dna: Uncovering universal latent forgery knowledge.arXiv preprint arXiv:2601.22515, 2026.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Dna: Uncovering universal latent forgery knowledge.arXiv preprint arXiv:2601.22515, 2026

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:49.155369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:49.155369Z digest=sha256:2a8af9899ae68a92f74c8871a37fb76ab0b37fde245b0383ce632908f3e1840d

Observation 7329596c-6e32-404e-8807-106b89078f3f · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Gemma: Open Models Based on Gemini Research and Technology

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:49.251487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:49.251487Z digest=sha256:f23abf64c5d433bdcfa740b3ef5d129cda4856f4557cd1e7d313cdcfe3028f1b

Observation afe3a815-d650-4634-8ae1-922834e6efcf · outbound

This paper cites FigStep: Jailbreaking large vision-language models via typographic visual prompts.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks FigStep: Jailbreaking large vision-language models via typographic visual prompts

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:59.736632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:49.356894Z digest=sha256:7a4b1fb239ba6b3fe0c81864aaae9e0fd51f56d7c7ef0d7322b9aa0bab995cf6

Observation a3904cd3-be16-40e6-ab81-a1f911ef0b4f · outbound

This paper cites The llama 3 herd of models, 2024.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks The llama 3 herd of models, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:49.592647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:49.592647Z digest=sha256:3541fc6793b0aed21ae002ea0b6727d6d4eb4eba4a54e1a5fb58774949795b8b

Observation 965dcc06-edae-409d-a29e-d0eabdd9e20e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:49.709052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:49.709052Z digest=sha256:301613132e256455eb24b7a8ec7b795bbddf4f22198a09b6f9aac2c6100b9933

Observation b3cac4d5-8985-4e00-974c-21b577d2d13d · outbound

This paper cites PKU-SafeRLHF: Towards multi- level safety alignment for LLMs with human preference.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks PKU-SafeRLHF: Towards multi- level safety alignment for LLMs with human preference

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:49.818291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:49.818291Z digest=sha256:63112a857dc5b5e68a01a90fd4aac3f2656ebeeceba0d5fd54c3e13d424a4d8a

Observation a5a78fc4-7c1b-4f4b-8762-7559e5ee21e5 · outbound

This paper cites an unresolved cited work.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:18:59.462162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:49.948433Z digest=sha256:2876b5f542b45d2fb4088b0f9021fe6baeb35a47397fe65372fc80ca741408cd

Observation 9efe5ece-fc1d-4b04-a61f-183b40f76cd2 · outbound

This paper cites Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, et al.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, et al

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:59.246967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:50.077899Z digest=sha256:0c0d1938698e4c92302e23220fbdf0a7ceacc59a977e252f79dc9e15cc1369b2

Observation 391afc04-b9e6-4eea-b279-73d6226f29f5 · outbound

This paper cites Images are achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Images are achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:59.008836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:50.254018Z digest=sha256:35f0df5115286078b395789cd01ebb903478cbf8c410f00b4511139cc68ec98a

Observation 95d81f17-dbd3-4b39-8784-51c43cd96727 · outbound

This paper cites TruthfulQA: Measuring how models mimic human falsehoods.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks TruthfulQA: Measuring how models mimic human falsehoods

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:58.773225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:50.386078Z digest=sha256:623f4e2d61d5f38208eea8e405b60effdd4e6f9e28432b81e35119dbb8edd12e

Observation b5f373f9-07a4-4730-bcde-6323a666b9bb · outbound

This paper cites Visual instruction tuning.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Visual instruction tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:50.496181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:50.496181Z digest=sha256:e363d2e7f2a218ba5577d327f4142e717cd8f279fd429a4c7fcdcd0d177dd958

Observation 642773eb-c65f-414d-ba00-0ccf0ba78eef · outbound

This paper cites MMBench: Is your multi-modal model an all- around player?, 2023.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks MMBench: Is your multi-modal model an all- around player?, 2023

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:50.613801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:50.613801Z digest=sha256:ef8ee89b910dde80fa1f153d268e0fab11d87995d6c811ec27a8d986162b5f49

Observation 517f68ac-7397-4396-b81b-ca04a9826e64 · outbound

This paper cites Decoupled weight decay regulariza- tion.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Decoupled weight decay regulariza- tion

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:50.721583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:50.721583Z digest=sha256:6e8c176a64deea0569e0afc47ec5b36dae6bf872a5c89e4a431a23fccec72f86

Observation 36e0cc30-eb5d-471d-84cc-5cd5421fb6a2 · outbound

This paper cites HarmBench: A standardized evaluation framework for automated red teaming and robust refusal.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks HarmBench: A standardized evaluation framework for automated red teaming and robust refusal

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:58.433158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:50.847333Z digest=sha256:258c1f985f87215889a046e9ad6ea468cdc2bf9d9affdf0c18efe3cde5c58103

Observation 83054a01-f0cd-4ff9-a024-7392f6604c96 · outbound

This paper cites Pruning convolutional neural networks for resource efficient inference.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Pruning convolutional neural networks for resource efficient inference

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:58.072871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:50.993536Z digest=sha256:00955fa6a8ec195401236f651b9d931e026e9ae3f75573bcd05a55f4d7417f16

Observation 777b2304-c637-43ff-89cf-f44664ae95d7 · outbound

This paper cites Training language models to follow instructions with human feedback.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Training language models to follow instructions with human feedback

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:57.645617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:51.108876Z digest=sha256:55e4a1468f969a2bafd3b20788def42f282b02949af1ed55b1d8abee917bde46

Observation 67ad9ef2-ede9-4589-83c2-ac872a61d6fc · outbound

This paper cites Representation noising: A defence mech- anism against harmful finetuning.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Representation noising: A defence mech- anism against harmful finetuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:57.275855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:51.280048Z digest=sha256:03fe68dde0d314b4c02ca353c0018aad085d0c6e470893dd1d366beb260d8657

Observation 3874d651-a443-4ef3-a1ff-56db695739b6 · outbound

This paper cites XSTest: A test suite for identifying exaggerated safety behaviours in large language models.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks XSTest: A test suite for identifying exaggerated safety behaviours in large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:56.952210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:51.435864Z digest=sha256:51b9f2b32703641785161f1babe5dafe0ca349e124e813865f56f11bb355675b

Observation b45e36db-5e0f-406e-a599-c1632c39a1f4 · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:51.590300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:51.590300Z digest=sha256:2e48dde63797a78be716391cc4496e359de5f5dbe1e1a20307fcbfebe72d0d3e

Observation 0382d27c-9de5-4212-a464-d3152fe14255 · outbound

This paper cites Where culture fades: revealing the cultural gap in text-to-image generation.arXiv preprint arXiv:2511.17282, 2025.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Where culture fades: revealing the cultural gap in text-to-image generation.arXiv preprint arXiv:2511.17282, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:51.980020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:51.980020Z digest=sha256:907cf66ea125cef497b00383746f941ce234191563581d5dec9a9bd74a70e91e

Observation f1ad0204-b8c7-497d-be98-9ce6e5cde755 · outbound

This paper cites TraceRouter: Robust safety for large foundation models via path-level intervention.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks TraceRouter: Robust safety for large foundation models via path-level intervention

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:56.579897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:52.226014Z digest=sha256:24538a54a1f9ff50ca16c13f66634d9886f1410c4aaa85effab45fdf13dbc8f6

Observation c64ee951-f054-4f4e-8519-3ab80e966c19 · outbound

This paper cites Orthoeraser: coupled-neuron orthogonal projection for concept erasure.arXiv preprint arXiv:2603.11493, 2026.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Orthoeraser: coupled-neuron orthogonal projection for concept erasure.arXiv preprint arXiv:2603.11493, 2026

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:52.311850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:52.311850Z digest=sha256:b4a819f3ad2c332d5a1d34be2411bc0af3f5b596860630970de2bbf3160d56f4

Observation 51dcc1f5-2ec3-4bba-abf7-f0e651ef2a17 · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks A StrongREJECT for Empty Jailbreaks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:52.404593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:52.404593Z digest=sha256:a42f89f4d782b32ed0680fb74f8e02a8ff5722da1b67366dd40671e06b68b6de

Observation 7018b5b6-ffcc-434f-8f23-acb594ded70c · outbound

This paper cites Dropout: A simple way to prevent neural networks from overfitting.Journal of Machine Learning Research, 15(56):1929–1958, 2014.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Dropout: A simple way to prevent neural networks from overfitting.Journal of Machine Learning Research, 15(56):1929–1958, 2014

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:52.471514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:52.471514Z digest=sha256:4ec299142cd5f08e99140f00ca75dd2b13767c1f95b3e9c7733cd24233ad9247

Observation 77401e9b-678a-40dc-9eaa-01507ef87998 · outbound

This paper cites Zico Kolter.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Zico Kolter

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:52.559100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:52.559100Z digest=sha256:829f05ca00ad9bcf8c0a4a245c4a4fb1e39c6a9f692ddb89c5d6760d2a0ee35b

Observation 3844070b-028c-4aed-b328-806cd2808c4f · outbound

This paper cites Zico Kolter.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Zico Kolter

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:56.252598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:52.637182Z digest=sha256:279db36b94a023da71959e5a2820efb1f4c1f4a530baa06703078d3ae62d3dd7

Observation 5881ff72-767d-43ea-ad4a-d7f368a58b0b · outbound

This paper cites SafeNeuron: Neuron- level safety alignment for large language models.arXiv preprint arXiv:2602.12158, 2026.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks SafeNeuron: Neuron- level safety alignment for large language models.arXiv preprint arXiv:2602.12158, 2026

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:52.703131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:52.703131Z digest=sha256:1eef625730283b2a1bbe6457ea95941859d32139729bae575ac79237516f511f

Observation 4d07861b-3234-4a57-b119-647180c55b74 · outbound

This paper cites Taxonomy, opportunities, and challenges of representation engi- neering for large language models.arXiv preprint arXiv:2502.19649, 2025.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Taxonomy, opportunities, and challenges of representation engi- neering for large language models.arXiv preprint arXiv:2502.19649, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:52.787762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:52.787762Z digest=sha256:cfcffafdba2a16e3d2cc0af1685b8b77695284be41fda34328aa5f660bf2d945

Observation b650a3e8-a2df-4c1a-b276-41f8ed9d33ba · outbound

This paper cites Jailbroken: How does LLM safety training fail? InAdvances in Neural Information Processing Systems, volume 36, 2023.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Jailbroken: How does LLM safety training fail? InAdvances in Neural Information Processing Systems, volume 36, 2023

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:55.871548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:52.858768Z digest=sha256:ec06641f77eb2742e8f6f262dad158b0ea2116f1bb6981e63bd055fd021a9990

Observation 4f1293c3-bd93-43c6-a8d7-123bee059ef4 · outbound

This paper cites NeuroStrike: Neuron-level attacks on aligned LLMs.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks NeuroStrike: Neuron-level attacks on aligned LLMs

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:55.461925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:52.926634Z digest=sha256:c0a0bca3b6633d15a343dbdb696b43d0274aed1a2ed91059dac77824ea8d1223

Observation 11b3413a-478c-4fd6-9f1d-e45d8ad267b3 · outbound

This paper cites NLSR: Neuron-level safety realignment of large language 9 models against harmful fine-tuning.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks NLSR: Neuron-level safety realignment of large language 9 models against harmful fine-tuning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:55.100559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:52.993493Z digest=sha256:5ba10cdaa3842142aba7a812f7c11799a3c0d110a28770e2d7b9a22c80a19320

Observation b4639a0e-6f5b-4dbe-8f57-3ad4a739267c · outbound

This paper cites Representation bending for large language model safety.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Representation bending for large language model safety

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:54.924763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:53.071474Z digest=sha256:4891295b26b2cb46f1cb245529fc18751493dbe9c2bc7230b2bde1b9601eec35

Observation 433c468d-dbc8-42b1-8e23-0e716e9f8166 · outbound

This paper cites Weston, and Xian Li.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Weston, and Xian Li

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:54.675581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:53.303482Z digest=sha256:0e51dbac52b4b654d7803a329388d8fce485cc556fe177ec929f72ceca07265c

Observation dc92f761-cf4f-40ed-91ff-dfb68f534e9c · outbound

This paper cites Understanding and enhancing safety mechanisms of LLMs via safety-specific neuron.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Understanding and enhancing safety mechanisms of LLMs via safety-specific neuron

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:54.473399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T00:18:53.415568Z digest=sha256:52835954ff27b9b1245469a8dca2b15f57039423d59e56e68e0d55eba54ace73

Observation 552e4e83-f8bc-4208-a42c-770f591127eb · outbound

This paper cites Improving Alignment and Robustness with Circuit Breakers.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Improving Alignment and Robustness with Circuit Breakers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:53.554407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:53.554407Z digest=sha256:bc2119fe8c918f4754c15fc5f0ea18fbcbb7ada31cf45cd4c5f211dcafddabb2

Observation c28e9822-e015-46e2-99c6-67dbb638237f · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:53.661855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:53.661855Z digest=sha256:d654a5f64b5d4f19cd988f8147c85bffd3326fff6e08b1a9486c02037b44720b

Pith citing papers

No inbound Pith citation observations are available.