Pith. sign in

Paper Citation Record · LEDGER

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety

As of 13 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2506.05451.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05451 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:24:26.477946Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T05:45:42.313410Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact2
  • verified fuzzy10
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dbf922d8-bb63-4726-ace1-fb395551418c · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Refusal in Language Models Is Mediated by a Single Direction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.305781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.305781Z digest=sha256:16b7e4791edc7c36afb757119d9de99feb79b6b5d70b866f380e10fdff71c6d6

Observation 062a7351-50e3-44b0-a0e1-ff7da1d591b2 · outbound

This paper cites Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.309978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.309978Z digest=sha256:b4cc369a3be3e275c4341fb99423d950a001f162d4b46e7f17e067cb610cfeac

Observation b58ac760-0795-4097-8005-5aaff2c6130e · outbound

This paper cites Discovering Latent Knowledge in Language Models Without Supervision.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Discovering Latent Knowledge in Language Models Without Supervision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.314508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.314508Z digest=sha256:b122bd4ee57c6d1e32bb96075bf78b46b65685117c476262f971b17a0b90b546

Observation a510d602-2812-4b6e-ad08-998edc02fac3 · outbound

This paper cites V ojtech Cahlik, Rodrigo Alves, and Pavel Kordik.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety V ojtech Cahlik, Rodrigo Alves, and Pavel Kordik

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.266765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:24:26.318627Z digest=sha256:eef39e2ffb0ba443128dc18c55921e88747e6ad2dd7892c3d4fd4a368bcec517

Observation e97702f2-a65c-42fc-8bb2-dd2a8b388c03 · outbound

This paper cites Nitay Calderon and Roi Reichart.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Nitay Calderon and Roi Reichart

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-07T10:24:27.077942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:24:26.323532Z digest=sha256:cd4d7901eda139481530d0ec88ad80303d364f94ab6fd054585cae1f193889a1

Observation d8ea653f-a2ff-40a9-9f0f-2877412ba779 · outbound

This paper cites Improving Steering Vectors by Targeting Sparse Autoencoder Features.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Improving Steering Vectors by Targeting Sparse Autoencoder Features

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.327706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.327706Z digest=sha256:5aaeafd162445b1c70740ac8ac6b973e9f1d922d536fb0a37c412a0fdede5e79

Observation c9e87b34-a8e6-4d4d-a985-b455905c92b1 · outbound

This paper cites Finetuning Language Models to Emit Linguistic Expressions of Uncertainty.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Finetuning Language Models to Emit Linguistic Expressions of Uncertainty

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.332694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.332694Z digest=sha256:4cec62af9ea4079f5e79737547e6e6ec139c3460e078bbfd4a36a6f7e6694fdf

Observation a08dd8f6-c46d-439f-8826-c1fd52b48c8c · outbound

This paper cites Faithful Reasoning Using Large Language Models.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Faithful Reasoning Using Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.340928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.340928Z digest=sha256:8b47bc01507eac4cc147a69992e925c1257c2800394ba55e888b295b8893844e

Observation f977c197-fd2d-41d8-aa68-130cebbea1dc · outbound

This paper cites InThe Eleventh International Conference on Learning Rep- resentations.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InThe Eleventh International Conference on Learning Rep- resentations

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.244661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:24:26.345165Z digest=sha256:0beaa67ca8944bf044ca3170cd7923fab7384d2f3971758d04f7cccaee8b9d02

Observation 07d74fd0-9172-4674-a7fe-6404471bb4ba · outbound

This paper cites Discovering Variable Binding Circuitry with Desiderata.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Discovering Variable Binding Circuitry with Desiderata

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.349316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.349316Z digest=sha256:723156d9ac145cfdaece5be61374db93afa7a792c990e37d4fbc0f44a3c222eb

Observation c72c0a66-3748-4fb3-b5e9-3db3bf864d17 · outbound

This paper cites Studying Large Language Model Generalization with Influence Functions.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Studying Large Language Model Generalization with Influence Functions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.357474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.357474Z digest=sha256:57ebbcb8467bdc9b1fa5da93918e34725e4346164b34d9f3b1c7b50825c5b05d

Observation 2711a5e4-750b-4a0e-abe1-ae7a0bf9a485 · outbound

This paper cites InProceedings of the 37th Interna- tional Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY , USA.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InProceedings of the 37th Interna- tional Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY , USA

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.220441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:24:26.366075Z digest=sha256:8fc6e1791ecf3b79141c1fec52580dc031e14f43243c55cc192dd2b63a1c49ed

Observation 55a3048b-7a0b-43f1-926f-4d7fd1a3dd6c · outbound

This paper cites Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.370726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.370726Z digest=sha256:5a4ce08360721c00a6aa5ef2ea729918287a8d1fe09d5288d728fb03d9eb7b37

Observation d010205d-0dd2-4b85-9ba4-9d27f86083e7 · outbound

This paper cites Zhengfu He, Xuyang Ge, Qiong Tang, Tianxiang Sun, Qinyuan Cheng, and Xipeng Qiu.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Zhengfu He, Xuyang Ge, Qiong Tang, Tianxiang Sun, Qinyuan Cheng, and Xipeng Qiu

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.375257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.375257Z digest=sha256:b6b441f30b6c07d8755cd54d3da0451965f6fe2289c4b7d463f8ab596223b20c

Observation 3488b3e3-dca5-4d5e-8340-3e9f76278872 · outbound

This paper cites Alon Jacovi and Yoav Goldberg.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Alon Jacovi and Yoav Goldberg

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.380068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.380068Z digest=sha256:491aea57cc5919c3dcf414a18eb19679e097855c6b84afa3976998e5959293e2

Observation d86483a8-da5a-482b-b580-4a6b03accbec · outbound

This paper cites How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.384700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.384700Z digest=sha256:01342566118d939936aa9a110b08bcf5d4b883e7d3782f48307b7db46a496287

Observation 271a24eb-f94f-46b3-b5ae-aaff77699d9a · outbound

This paper cites InProceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY , USA.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InProceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY , USA

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.209641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:24:26.392317Z digest=sha256:b4718673ae2cd17f143e9394234e98c1b03e9b6e005ad8dc46559f12f523aa6f

Observation 237d5644-9cb1-4701-bc63-1347536b7e4c · outbound

This paper cites InProceedings of the 61st Annual Meeting of the Association for Compu- tational Linguistics (Volume 3: System Demonstra- tions), pages 42–50, Toronto, Canada.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InProceedings of the 61st Annual Meeting of the Association for Compu- tational Linguistics (Volume 3: System Demonstra- tions), pages 42–50, Toronto, Canada

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.196551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:24:26.396707Z digest=sha256:e0c8cbf9475d447405b7cce7f885d9732e1a4eab972212df7f626a865bf8cafa

Observation d9c0bf9b-88c6-4aae-b144-8b6267082701 · outbound

This paper cites InThe Twelfth International Conference on Learning Repre- sentations.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InThe Twelfth International Conference on Learning Repre- sentations

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.184327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:24:26.400384Z digest=sha256:64bcfe2e4300f1fa466b6837b2ec2cb2f699816db261eb8e37370947b7d2b81c

Observation 6532def3-7402-47d1-89d6-aeee8c336770 · outbound

This paper cites Rethinking Explainability as a Dialogue: A Practitioner's Perspective.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Rethinking Explainability as a Dialogue: A Practitioner's Perspective

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.403709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.403709Z digest=sha256:8926b1f589becbe89221d39fbe3ef42d6ac6dc8901c986f6ca6232b1daa62b46

Observation f0a5476d-6ef2-4063-9cb9-57abc06527db · outbound

This paper cites HuDEx: Integrating Hallucination Detection and Explainability for Enhancing the Reliability of LLM responses.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety HuDEx: Integrating Hallucination Detection and Explainability for Enhancing the Reliability of LLM responses

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:24:26.766381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:24:26.408324Z digest=sha256:71686fcd2d0c6456e96ebae5ce9102e17e24fb58f86ae22619f7b2f0e7973335

Observation 9d704216-f085-43b4-ad3d-bf434267d1ac · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.412268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.412268Z digest=sha256:39acc36922b9d662668898777135439206b03b8f9f23c4d21f84cf9b2ce0a929

Observation d9ef54eb-60ab-41e0-90dd-9dd62454179e · outbound

This paper cites Interpretable-by-Design Text Understanding with Iteratively Generated Concept Bottleneck.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Interpretable-by-Design Text Understanding with Iteratively Generated Concept Bottleneck

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.415954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.415954Z digest=sha256:c5149108c5fce2bb7dcf721b67104b01b1271cd0f0ae8b860380d25d05496441

Observation b7f2c7cf-d0a3-41ba-bb9f-906cba85e495 · outbound

This paper cites Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.419920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.419920Z digest=sha256:8bbe49412e6f787d8348540fb06578c9ee5779b11a54ed456e74185d203e777f

Observation de2e084d-6a75-4237-9fca-33dbbf571323 · outbound

This paper cites SaRO: Enhancing LLM Safety through Reasoning-based Alignment.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety SaRO: Enhancing LLM Safety through Reasoning-based Alignment

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.427788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.427788Z digest=sha256:f2793e6b7338a4b86ad432bcf9fc5bb82aecbaa7a8290e8bbe403056e1a14cf1

Observation 686d0779-98ad-4775-9f85-46cf60b7591e · outbound

This paper cites Gradient-Based Automated Iterative Recovery for Parameter-Efficient Tuning.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Gradient-Based Automated Iterative Recovery for Parameter-Efficient Tuning

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:24:26.697247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:24:26.431523Z digest=sha256:2ab97d69176e8fffea44033678c18b179514db5c343f167de3f945fa8eb9fd33

Observation cbeb1e1e-31e7-498f-af5f-ca4aa00efc3b · outbound

This paper cites Steering Language Model Refusal with Sparse Autoencoders.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Steering Language Model Refusal with Sparse Autoencoders

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.435845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.435845Z digest=sha256:e1982d7e46d22fe5c518ba02f1f4d4523bc894e9ca564820e33224384f63f3f0

Observation 08cddfeb-c7d8-4643-b109-c0b071d77632 · outbound

This paper cites Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.440712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.440712Z digest=sha256:0b9f936c4991466cb4ed3251830aa4b28d48f01b17b63ab49d97c964e04a6abe

Observation d2f08816-1d81-41f5-87dd-077d24031298 · outbound

This paper cites RevPRAG: Revealing Poisoning Attacks in Retrieval-Augmented Generation through LLM Activation Analysis.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety RevPRAG: Revealing Poisoning Attacks in Retrieval-Augmented Generation through LLM Activation Analysis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.452708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.452708Z digest=sha256:03863580a31463281e3b4f4afba7076f27b13a615966400a71f914b1438f0e41

Observation 25099a16-f626-442c-8027-75aa87f62c6f · outbound

This paper cites InICLR 2025 Workshop on Building Trust in Language Models and Applications.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InICLR 2025 Workshop on Building Trust in Language Models and Applications

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.160267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:24:26.460867Z digest=sha256:b30deb9d9142abeefe83069f749bb2c8b4c3715c25bc9cb509425a9f4f244e9a

Observation 18f8d42c-9149-455b-a29f-653f71b8df80 · outbound

This paper cites Steering Language Models With Activation Engineering.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Steering Language Models With Activation Engineering

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.465201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.465201Z digest=sha256:b1db365dece7356e58a9f3b04fc165ae3eb52c0b068bca919a25366b88e0b73c

Observation b7093f4e-d86c-4039-9e92-9ea347a812a3 · outbound

This paper cites an unresolved cited work.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:24:27.148900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:24:26.469980Z digest=sha256:726b949e691d5e39ff62c7477bc51f71228deda088ec04183d3f4b66573da168

Observation 6077e1dc-c411-4eb3-8de1-7e23de6a6069 · outbound

This paper cites Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.474285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.474285Z digest=sha256:447e2fa0f986e9c8c550e21b1a81f39a0dc41aed26a250629488dad624682320

Observation d632adf3-891d-4c43-9da3-7e8ddee5ac65 · outbound

This paper cites Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.477946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.477946Z digest=sha256:28b2f592753d905b1f39dcd86eb9afdff0e38454d889a9eb43340c1e3ba5308b

Observation fb2c1e6d-62a1-48f1-99ca-61df1f20c819 · outbound

This paper cites why should i trust you?.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety why should i trust you?

Reference 483

Resolution
verified exact
raw_fallback, observed 2026-08-07T10:24:26.660635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:24:26.444527Z digest=sha256:e928c11af92f9a67537f45b3735edda49f6e08220fb7f9025e36855aec408e92

Observation 99d8ab01-3987-47fe-a284-cd5793d76918 · outbound

This paper cites Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.423582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.423582Z digest=sha256:3e83f156750feebc5a3782f024900dd25312d0b2b020d3c2e93647710da8aebb

Observation 55ba416a-3575-4f8b-982a-9f11b3f56127 · outbound

This paper cites What do you learn from context? Probing for sentence structure in contextualized word representations.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety What do you learn from context? Probing for sentence structure in contextualized word representations

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.456815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.456815Z digest=sha256:4247726bfa6a2d7b05988c74211b43fe6c3efb30e41713ce101d864fa1e3a331

Observation d9acd6c1-29b8-4742-8f4c-bda55b6e9ef2 · outbound

This paper cites In Proceedings of the 58th Annual Meeting of the Asso- ciation for Computational Linguistics, pages 5553– 5563, Online.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety In Proceedings of the 58th Annual Meeting of the Asso- ciation for Computational Linguistics, pages 5553– 5563, Online

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.232109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:24:26.362341Z digest=sha256:a1e46e1b44c412a51a0bc9577729ffa12bc2274ba97b759b90c4554fd8c31256

Observation 2848e556-47e8-4bc3-8182-b773c9fa1e82 · outbound

This paper cites Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.353163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.353163Z digest=sha256:7ba9eb6f550412da4980603359d1dd6c46dc7320817dd555f4af1be376de053f

Observation 1215c80c-f397-4304-89cd-3a2e21d794a6 · outbound

This paper cites AtP*: An efficient and scalable method for localizing LLM behaviour to components.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety AtP*: An efficient and scalable method for localizing LLM behaviour to components

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.388763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.388763Z digest=sha256:c2d0237b5c022dade7ded75b4663054589952a358a9f5288c350c73fe6f9f7c6

Observation 2f1c746e-ede6-4330-bb77-3f4bf6832ede · outbound

This paper cites InAdvances in Neural Infor- mation Processing Systems, volume 36, pages 16318– 16352.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InAdvances in Neural Infor- mation Processing Systems, volume 36, pages 16318– 16352

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.255379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:24:26.336515Z digest=sha256:0300054fa11f1393073143d8699cb12880de0e1247db8710fd2211466bf48ec6

Observation ebd8f26f-5dc9-4d52-9ac6-bef8b3335c81 · outbound

This paper cites Get my drift? Catching LLM Task Drift with Activation Deltas.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Get my drift? Catching LLM Task Drift with Activation Deltas

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.300328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.300328Z digest=sha256:7eeb114930d421cf6a1eb1d41825b2346fe816d630df03b8ae4d3a5ecd13b264

Observation 560521ff-a5b7-43da-a648-05155e0ef7b0 · outbound

This paper cites SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs

Reference 2025

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:24:27.136657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:24:26.294848Z digest=sha256:bb3db5cdbf55293c69b6c470e4e3164b306f0b040d596658a6380ee684143bc9

Observation 9e69ad70-b6de-42ab-8d6e-6831fa30682f · outbound

This paper cites Aaquib Syed, Can Rager, and Arthur Conmy.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Aaquib Syed, Can Rager, and Arthur Conmy

Reference 3328

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.172301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:24:26.448608Z digest=sha256:475d303dca5975f1a103f00beaf9f79eab4dcd7925e0459b47bb10ee4727c32a

Pith citing papers

Observation c0ba13cc-7577-4cd3-9ca2-2dfa84908394 · inbound

Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models cites this paper.

Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-03T05:45:42.313410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:45:42.313410Z digest=sha256:066e9ba9f2014a864f230092d1b44fe3139fdbe05b1f8425bd45815d740d3219