Pith. sign in

Paper Citation Record · LEDGER

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety

As of 14 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2506.05451.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05451 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:24:26.477946Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T05:45:42.313410Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact2
  • verified fuzzy10
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dbf922d8-bb63-4726-ace1-fb395551418c · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Refusal in Language Models Is Mediated by a Single Direction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.305781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.305781Z digest=sha256:16b7e4791edc7c36afb757119d9de99feb79b6b5d70b866f380e10fdff71c6d6

Observation 062a7351-50e3-44b0-a0e1-ff7da1d591b2 · outbound

This paper cites Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.309978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.309978Z digest=sha256:b4cc369a3be3e275c4341fb99423d950a001f162d4b46e7f17e067cb610cfeac

Observation b58ac760-0795-4097-8005-5aaff2c6130e · outbound

This paper cites Discovering Latent Knowledge in Language Models Without Supervision.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Discovering Latent Knowledge in Language Models Without Supervision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.314508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.314508Z digest=sha256:b122bd4ee57c6d1e32bb96075bf78b46b65685117c476262f971b17a0b90b546

Observation a510d602-2812-4b6e-ad08-998edc02fac3 · outbound

This paper cites V ojtech Cahlik, Rodrigo Alves, and Pavel Kordik.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety V ojtech Cahlik, Rodrigo Alves, and Pavel Kordik

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.266765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:24:26.318627Z digest=sha256:525b08fc030e8ce2725103b7641b233139128799f8439c37ba069d03a05637c7

Observation e97702f2-a65c-42fc-8bb2-dd2a8b388c03 · outbound

This paper cites Nitay Calderon and Roi Reichart.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Nitay Calderon and Roi Reichart

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-07T10:24:27.077942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:24:26.323532Z digest=sha256:49b9c4814a0eb8b8787779e2b649ba08bfb8ccefbbe7b0a56f702be49993773f

Observation d8ea653f-a2ff-40a9-9f0f-2877412ba779 · outbound

This paper cites Improving Steering Vectors by Targeting Sparse Autoencoder Features.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Improving Steering Vectors by Targeting Sparse Autoencoder Features

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.327706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.327706Z digest=sha256:5aaeafd162445b1c70740ac8ac6b973e9f1d922d536fb0a37c412a0fdede5e79

Observation c9e87b34-a8e6-4d4d-a985-b455905c92b1 · outbound

This paper cites Finetuning Language Models to Emit Linguistic Expressions of Uncertainty.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Finetuning Language Models to Emit Linguistic Expressions of Uncertainty

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.332694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.332694Z digest=sha256:4cec62af9ea4079f5e79737547e6e6ec139c3460e078bbfd4a36a6f7e6694fdf

Observation a08dd8f6-c46d-439f-8826-c1fd52b48c8c · outbound

This paper cites Faithful Reasoning Using Large Language Models.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Faithful Reasoning Using Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.340928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.340928Z digest=sha256:8b47bc01507eac4cc147a69992e925c1257c2800394ba55e888b295b8893844e

Observation f977c197-fd2d-41d8-aa68-130cebbea1dc · outbound

This paper cites InThe Eleventh International Conference on Learning Rep- resentations.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InThe Eleventh International Conference on Learning Rep- resentations

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.244661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:24:26.345165Z digest=sha256:c56c1439f8a2be8016817bd7064e73b61f89b1945d26a0222bfa7cf82290b122

Observation 07d74fd0-9172-4674-a7fe-6404471bb4ba · outbound

This paper cites Discovering Variable Binding Circuitry with Desiderata.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Discovering Variable Binding Circuitry with Desiderata

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.349316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.349316Z digest=sha256:723156d9ac145cfdaece5be61374db93afa7a792c990e37d4fbc0f44a3c222eb

Observation c72c0a66-3748-4fb3-b5e9-3db3bf864d17 · outbound

This paper cites Studying Large Language Model Generalization with Influence Functions.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Studying Large Language Model Generalization with Influence Functions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.357474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.357474Z digest=sha256:57ebbcb8467bdc9b1fa5da93918e34725e4346164b34d9f3b1c7b50825c5b05d

Observation 2711a5e4-750b-4a0e-abe1-ae7a0bf9a485 · outbound

This paper cites InProceedings of the 37th Interna- tional Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY , USA.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InProceedings of the 37th Interna- tional Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY , USA

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.220441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:24:26.366075Z digest=sha256:f611d3e8dc3b45642659b2b44de9e384a94697395d5c1c214dd3024a533487e0

Observation 55a3048b-7a0b-43f1-926f-4d7fd1a3dd6c · outbound

This paper cites Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.370726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.370726Z digest=sha256:5a4ce08360721c00a6aa5ef2ea729918287a8d1fe09d5288d728fb03d9eb7b37

Observation d010205d-0dd2-4b85-9ba4-9d27f86083e7 · outbound

This paper cites Zhengfu He, Xuyang Ge, Qiong Tang, Tianxiang Sun, Qinyuan Cheng, and Xipeng Qiu.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Zhengfu He, Xuyang Ge, Qiong Tang, Tianxiang Sun, Qinyuan Cheng, and Xipeng Qiu

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.375257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.375257Z digest=sha256:b6b441f30b6c07d8755cd54d3da0451965f6fe2289c4b7d463f8ab596223b20c

Observation 3488b3e3-dca5-4d5e-8340-3e9f76278872 · outbound

This paper cites Alon Jacovi and Yoav Goldberg.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Alon Jacovi and Yoav Goldberg

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.380068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.380068Z digest=sha256:491aea57cc5919c3dcf414a18eb19679e097855c6b84afa3976998e5959293e2

Observation d86483a8-da5a-482b-b580-4a6b03accbec · outbound

This paper cites How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.384700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.384700Z digest=sha256:01342566118d939936aa9a110b08bcf5d4b883e7d3782f48307b7db46a496287

Observation 271a24eb-f94f-46b3-b5ae-aaff77699d9a · outbound

This paper cites InProceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY , USA.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InProceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY , USA

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.209641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:24:26.392317Z digest=sha256:38d8996d840ac2dfe5a45a873c9cd06483eafb9ced5a4e520245bc67f7d7039c

Observation 237d5644-9cb1-4701-bc63-1347536b7e4c · outbound

This paper cites InProceedings of the 61st Annual Meeting of the Association for Compu- tational Linguistics (Volume 3: System Demonstra- tions), pages 42–50, Toronto, Canada.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InProceedings of the 61st Annual Meeting of the Association for Compu- tational Linguistics (Volume 3: System Demonstra- tions), pages 42–50, Toronto, Canada

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.196551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:24:26.396707Z digest=sha256:67bf1796b06813f653590fb5ca10c472130883d6b81331a8fce9cb9705cd9643

Observation d9c0bf9b-88c6-4aae-b144-8b6267082701 · outbound

This paper cites InThe Twelfth International Conference on Learning Repre- sentations.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InThe Twelfth International Conference on Learning Repre- sentations

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.184327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:24:26.400384Z digest=sha256:750468925f9de468b3094eb8cb5b4bec248ce1deda85b7049ac84a9ea640d023

Observation 6532def3-7402-47d1-89d6-aeee8c336770 · outbound

This paper cites Rethinking Explainability as a Dialogue: A Practitioner's Perspective.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Rethinking Explainability as a Dialogue: A Practitioner's Perspective

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.403709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.403709Z digest=sha256:c508148db3923afc076ed13259382d711f1ec505c0622e92a12cc064ef01972d

Observation f0a5476d-6ef2-4063-9cb9-57abc06527db · outbound

This paper cites HuDEx: Integrating Hallucination Detection and Explainability for Enhancing the Reliability of LLM responses.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety HuDEx: Integrating Hallucination Detection and Explainability for Enhancing the Reliability of LLM responses

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:24:26.766381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:24:26.408324Z digest=sha256:a2ff6621af20f09285c0249f0088c1400740cc250266dd5fe84e1dfdb4eabfa5

Observation 9d704216-f085-43b4-ad3d-bf434267d1ac · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.412268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.412268Z digest=sha256:db1d42c434c57df69f0a1f0fcddba3bca330595aa0066c656404c22e3f161341

Observation d9ef54eb-60ab-41e0-90dd-9dd62454179e · outbound

This paper cites Interpretable-by-Design Text Understanding with Iteratively Generated Concept Bottleneck.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Interpretable-by-Design Text Understanding with Iteratively Generated Concept Bottleneck

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.415954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.415954Z digest=sha256:c5149108c5fce2bb7dcf721b67104b01b1271cd0f0ae8b860380d25d05496441

Observation b7f2c7cf-d0a3-41ba-bb9f-906cba85e495 · outbound

This paper cites Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.419920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.419920Z digest=sha256:8bbe49412e6f787d8348540fb06578c9ee5779b11a54ed456e74185d203e777f

Observation de2e084d-6a75-4237-9fca-33dbbf571323 · outbound

This paper cites SaRO: Enhancing LLM Safety through Reasoning-based Alignment.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety SaRO: Enhancing LLM Safety through Reasoning-based Alignment

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.427788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.427788Z digest=sha256:f2793e6b7338a4b86ad432bcf9fc5bb82aecbaa7a8290e8bbe403056e1a14cf1

Observation 686d0779-98ad-4775-9f85-46cf60b7591e · outbound

This paper cites Gradient-Based Automated Iterative Recovery for Parameter-Efficient Tuning.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Gradient-Based Automated Iterative Recovery for Parameter-Efficient Tuning

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:24:26.697247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:24:26.431523Z digest=sha256:610cea4996892e95cd128422da5d3db4fdbcbf83031fe23ab7bf0099234b5fcc

Observation cbeb1e1e-31e7-498f-af5f-ca4aa00efc3b · outbound

This paper cites Steering Language Model Refusal with Sparse Autoencoders.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Steering Language Model Refusal with Sparse Autoencoders

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.435845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.435845Z digest=sha256:3ed4c4dc91225c4860d9e081f15e2059acc8d47d7c6d849f2d0b798e94012f76

Observation 08cddfeb-c7d8-4643-b109-c0b071d77632 · outbound

This paper cites Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.440712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.440712Z digest=sha256:b1809db260044d5da20fd46382ce2bae68d9c28b3414ac108e7030019153541d

Observation d2f08816-1d81-41f5-87dd-077d24031298 · outbound

This paper cites RevPRAG: Revealing Poisoning Attacks in Retrieval-Augmented Generation through LLM Activation Analysis.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety RevPRAG: Revealing Poisoning Attacks in Retrieval-Augmented Generation through LLM Activation Analysis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.452708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.452708Z digest=sha256:03863580a31463281e3b4f4afba7076f27b13a615966400a71f914b1438f0e41

Observation 25099a16-f626-442c-8027-75aa87f62c6f · outbound

This paper cites InICLR 2025 Workshop on Building Trust in Language Models and Applications.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InICLR 2025 Workshop on Building Trust in Language Models and Applications

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.160267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:24:26.460867Z digest=sha256:8589f929e36526a16872db8a02b8a3eed040a835e5010aea80ccf70733a3815f

Observation 18f8d42c-9149-455b-a29f-653f71b8df80 · outbound

This paper cites Steering Language Models With Activation Engineering.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Steering Language Models With Activation Engineering

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.465201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.465201Z digest=sha256:b1db365dece7356e58a9f3b04fc165ae3eb52c0b068bca919a25366b88e0b73c

Observation b7093f4e-d86c-4039-9e92-9ea347a812a3 · outbound

This paper cites an unresolved cited work.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:24:27.148900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:24:26.469980Z digest=sha256:7ad2c4354b3487b06e6b000bc4fc8dac95d9989d1cc28eb784e1e4ee58235e45

Observation 6077e1dc-c411-4eb3-8de1-7e23de6a6069 · outbound

This paper cites Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.474285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.474285Z digest=sha256:447e2fa0f986e9c8c550e21b1a81f39a0dc41aed26a250629488dad624682320

Observation d632adf3-891d-4c43-9da3-7e8ddee5ac65 · outbound

This paper cites Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.477946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.477946Z digest=sha256:28b2f592753d905b1f39dcd86eb9afdff0e38454d889a9eb43340c1e3ba5308b

Observation fb2c1e6d-62a1-48f1-99ca-61df1f20c819 · outbound

This paper cites why should i trust you?.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety why should i trust you?

Reference 483

Resolution
verified exact
raw_fallback, observed 2026-08-07T10:24:26.660635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:24:26.444527Z digest=sha256:e34722e65fc56a2be13abf790ae83e120ef476d738d19bd3a589988899a70374

Observation 99d8ab01-3987-47fe-a284-cd5793d76918 · outbound

This paper cites Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.423582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.423582Z digest=sha256:3e83f156750feebc5a3782f024900dd25312d0b2b020d3c2e93647710da8aebb

Observation 55ba416a-3575-4f8b-982a-9f11b3f56127 · outbound

This paper cites What do you learn from context? Probing for sentence structure in contextualized word representations.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety What do you learn from context? Probing for sentence structure in contextualized word representations

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.456815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.456815Z digest=sha256:4247726bfa6a2d7b05988c74211b43fe6c3efb30e41713ce101d864fa1e3a331

Observation d9acd6c1-29b8-4742-8f4c-bda55b6e9ef2 · outbound

This paper cites In Proceedings of the 58th Annual Meeting of the Asso- ciation for Computational Linguistics, pages 5553– 5563, Online.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety In Proceedings of the 58th Annual Meeting of the Asso- ciation for Computational Linguistics, pages 5553– 5563, Online

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.232109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:24:26.362341Z digest=sha256:f53984480f5a384580cdfd88162bf6b7475831ccfca017a4f9a09a3a563f67e9

Observation 2848e556-47e8-4bc3-8182-b773c9fa1e82 · outbound

This paper cites Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.353163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.353163Z digest=sha256:7ba9eb6f550412da4980603359d1dd6c46dc7320817dd555f4af1be376de053f

Observation 1215c80c-f397-4304-89cd-3a2e21d794a6 · outbound

This paper cites AtP*: An efficient and scalable method for localizing LLM behaviour to components.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety AtP*: An efficient and scalable method for localizing LLM behaviour to components

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.388763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.388763Z digest=sha256:c2d0237b5c022dade7ded75b4663054589952a358a9f5288c350c73fe6f9f7c6

Observation 2f1c746e-ede6-4330-bb77-3f4bf6832ede · outbound

This paper cites InAdvances in Neural Infor- mation Processing Systems, volume 36, pages 16318– 16352.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InAdvances in Neural Infor- mation Processing Systems, volume 36, pages 16318– 16352

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.255379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:24:26.336515Z digest=sha256:d4719032a627e52c1b8c38cfc21c9ad0c0763bd53eb50c60b70e845096fb6514

Observation ebd8f26f-5dc9-4d52-9ac6-bef8b3335c81 · outbound

This paper cites Get my drift? Catching LLM Task Drift with Activation Deltas.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Get my drift? Catching LLM Task Drift with Activation Deltas

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.300328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.300328Z digest=sha256:37bd42a7a7edf581a07d9af69970729893ad23ea73bda90b3ddc238ab8dc510d

Observation 560521ff-a5b7-43da-a648-05155e0ef7b0 · outbound

This paper cites SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs

Reference 2025

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:24:27.136657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:24:26.294848Z digest=sha256:59f04ed28aab6c86ca60ded5f478272ff67de58ffe3f8f8f9ad837e26fae13b3

Observation 9e69ad70-b6de-42ab-8d6e-6831fa30682f · outbound

This paper cites Aaquib Syed, Can Rager, and Arthur Conmy.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Aaquib Syed, Can Rager, and Arthur Conmy

Reference 3328

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.172301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T10:24:26.448608Z digest=sha256:d8fd49c3d124d57e856d6cb5f425346136bf125044888b9c5bd555e0246f21fb

Pith citing papers

Observation c0ba13cc-7577-4cd3-9ca2-2dfa84908394 · inbound

Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models cites this paper.

Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-03T05:45:42.313410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:45:42.313410Z digest=sha256:066e9ba9f2014a864f230092d1b44fe3139fdbe05b1f8425bd45815d740d3219