Pith. sign in

Paper Citation Record · LEDGER

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs

As of 23 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 2 inbound Pith citation observations for arXiv:2507.04365.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04365 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:54:12.847460Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:54:06.538826Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:26:59.311234Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed19b6ad-0d61-44f6-82e2-19592791cf44 · outbound

This paper cites Bowman, Ethan Perez, Roger B.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Bowman, Ethan Perez, Roger B

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:16.871531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:54:09.120032Z digest=sha256:1af245b044d514a99b1f9b0f115622b999ff73eff8479c67b114dd0be003c860

Observation 5b845d7d-de3a-4664-96bf-60899d8be88b · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:09.296540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:09.296540Z digest=sha256:2373bacfd323237c9d6d83872aaa0cf2e301aeeb81dc13f66e8367afcc73912c

Observation f66d6f79-7d2c-4b13-b5e1-5eb1f2785950 · outbound

This paper cites Zico Kolter.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Zico Kolter

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:16.525171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:54:09.403630Z digest=sha256:f125bd4bc24dc45ee9b14a4abf122a7f118f39d7a2ac01cd0976fe2aa9646d97

Observation d07ac786-cf88-49bb-96db-b3fbc6d406e1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:09.473833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:09.473833Z digest=sha256:72ed9b407aed7ce3cf3ea283e9ac0810739ebdb476dab0991b8f3f7e2a83ed67

Observation d46fb15d-4d9e-4bac-8e2d-91c6fcac36ae · outbound

This paper cites an unresolved cited work.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:54:16.254930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:54:09.551089Z digest=sha256:f2e29972fcc2be095d43612a0c57b3c15ec4476785460a2c0b3bfc31dd8a6878

Observation 48d1a98c-12c5-459b-af6c-feb27d1f27f5 · outbound

This paper cites Token highlighter: Inspecting and mitigating jailbreak prompts for large language models.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Token highlighter: Inspecting and mitigating jailbreak prompts for large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:16.027866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:54:09.611247Z digest=sha256:5b04907812a166c8f0acf664aba223829c4937bbea94871dfa60e46ae94318b4

Observation bf8fcf33-b58b-4596-9a6d-f3fd30712868 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:09.684400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:09.684400Z digest=sha256:e9b26ab909772a9831624d9374794c339f749386e1ebd7c73d3dae5eb7a5886b

Observation 87b3f59d-53ea-4897-9fb7-2616726aaa25 · outbound

This paper cites an unresolved cited work.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:09.839439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:09.839439Z digest=sha256:3432b18b728231e0a3a7cb72afb2ef7f57fe698ccf8eb79ee035945b43e26e2f

Observation 65e8a31d-8f11-4111-b10a-3e6f44fddd25 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:09.982978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:09.982978Z digest=sha256:7e7e6804372fcf66c7c0023472f89b110dff252d9abb365e8e8f74c0eca0b77c

Observation c10190c4-b752-494d-a3d4-1bc75b923637 · outbound

This paper cites Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:10.094408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:10.094408Z digest=sha256:5b92e0b4f3c20bd29e57a8909a1f90e190b7c9c0e147268c6578c7c388711376

Observation 46b3415f-05b0-4731-a310-adfb9bc3f1ba · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:10.231324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:10.231324Z digest=sha256:ea89ba09fbba54bf441fb002ea125c6a4f7a4f9cbc579f278d3283318ec2b57a

Observation aa504a37-0b53-40b5-90e8-832619a453e6 · outbound

This paper cites GPT-4 Technical Report.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs GPT-4 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:10.352809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:10.352809Z digest=sha256:6e4c27dddfcccb813d6a476add3474b1853c5789000e97810cc97337fb5247f9

Observation 15698224-46b7-4e96-9c96-bffdb51d7ee5 · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:10.497216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:10.497216Z digest=sha256:b2fe5577438c413e520c0295fc69be11e57f968deb78227d01c408d41db84a35

Observation 803559cc-ce6a-414a-801e-0e4e2297a345 · outbound

This paper cites an unresolved cited work.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:54:15.761278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:54:10.684287Z digest=sha256:f878ba2fd196ec9f1d122eb19fb770e3468b4bef087ce5db2b1d1d83a4d9b066

Observation a7a98138-4888-49b4-919a-bd33944a6e4b · outbound

This paper cites Qwen2.5 Technical Report.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Qwen2.5 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:10.804884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:10.804884Z digest=sha256:b901a16aae21179f861aafaed3af269069b4587579d3877f8691e9541d7b5581

Observation 34c8d635-0166-48a9-8601-43fe20c4b908 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:15.522650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:54:10.923995Z digest=sha256:179c4fe14344867ee856387f275fdd527949617c02c9a83742899d3b9d6fc73a

Observation ed5bec15-6373-4d12-a047-fdc5659bfdd5 · outbound

This paper cites Yu, Qingsong Wen, and Yang Liu.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Yu, Qingsong Wen, and Yang Liu

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:15.229719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:54:11.081110Z digest=sha256:8500ad38eb47561576372ab8200711c371aef63ef1e3ae375e13b2b42dbb74b4

Observation 5b8c7c9a-59fd-4c5b-854f-6071827ddde2 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Jailbroken: How Does LLM Safety Training Fail?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:11.189415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:11.189415Z digest=sha256:58d326f8d9dbe0407016c493e165d5576d9c7477aafa02e300df5b7759a4c899

Observation 281b06a9-2921-4fd4-9b5a-43addc867693 · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminders.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Defending chatgpt against jailbreak attack via self-reminders

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:14.987500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:54:11.364637Z digest=sha256:4435e458f56c377ada30484017e85d763603d6bab9009e6beecd8f8ddde212a7

Observation 44031e8f-993a-4c66-8791-99f1696146f0 · outbound

This paper cites Safedecoding: Defending against jailbreak attacks via safety-aware decoding.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Safedecoding: Defending against jailbreak attacks via safety-aware decoding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:14.695109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:54:11.530520Z digest=sha256:ea453b5fe90b0cc9311d9cccc865f452905842d758c75dfdf91516ccacb66d70

Observation 3d4ba01d-58e9-4a56-8fad-03f5838ddbae · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:11.686321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:11.686321Z digest=sha256:bd135d2fd15b90e3bee674f245e6250a31f54801a89049405005798f77be51b8

Observation e392ee3c-fb5c-481c-9615-e62c0bd6c656 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Low-Resource Languages Jailbreak GPT-4

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:11.798190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:11.798190Z digest=sha256:8de69148f2324744a93a057dbb14068df06ee7d918649b6b32944fef2f2d9165

Observation ba7b66aa-1fbd-4999-9ebd-7bdd9a34e7b0 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:11.959670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:11.959670Z digest=sha256:ca52228f56b1e1b72bc344722039e77eae4d23b91cb5362beae574257ace42af

Observation 046c14d0-ecf8-4438-a0aa-a7d59829a5b7 · outbound

This paper cites Each matrix has dimensions d × d, and there are four such matrices: Memory for attention matrices = 4d2.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Each matrix has dimensions d × d, and there are four such matrices: Memory for attention matrices = 4d2

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:14.418062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:54:12.083145Z digest=sha256:97671e57b36dba0c5aeb8e5c73e64337012aacef99c6de660f99df33dcdaa020

Observation 86198f14-e088-479d-9735-6c0ccc9a442a · outbound

This paper cites The first maps the input dimension d to an intermediate dimension 4d, and the second maps back to d.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs The first maps the input dimension d to an intermediate dimension 4d, and the second maps back to d

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:14.162079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:54:12.219701Z digest=sha256:31ed0cb654ffac0fc64fe8b7a62d15925942c39b664dd477175bfe5882dd8e63

Observation 8f7bb32d-8ba1-40fc-a6e7-876053682a09 · outbound

This paper cites With n + m tokens in total (e.g., n input tokens and m output tokens), the memory required for Keys and Values per layer is: Key/Value Memory per layer (bytes) = 4(n + m)d.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs With n + m tokens in total (e.g., n input tokens and m output tokens), the memory required for Keys and Values per layer is: Key/Value Memory per layer (bytes) = 4(n + m)d

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:13.924016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:54:12.366585Z digest=sha256:daf5914eb5582d216e24b73aee220cf80e84c346687e276aa849d1681646d729

Observation 234e7a8c-c05b-45d5-8c90-ef33670f17bb · outbound

This paper cites an unresolved cited work.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:54:13.681569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:54:12.524515Z digest=sha256:8cc256de898091de0561596555c8b98ca10189f181579281bc7f148f13e30b6f

Observation 50d55ebd-4441-4f3a-8d0a-1931f66db025 · outbound

This paper cites an unresolved cited work.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:54:13.410898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:54:12.695682Z digest=sha256:555d23c8385d13f414e188a607024c7c123c772d94f21546dd2e8cd3eff67f04

Observation e9d4c793-97d0-4547-a529-97ef0be96397 · outbound

This paper cites an unresolved cited work.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:54:13.209306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:54:12.847460Z digest=sha256:624b9f6b88e62d40809a34f17b76780219fc439af734fe6b8e4cea2779e6ddf0

Pith citing papers

Observation 5f1b3b36-483b-4781-ada2-c313b1e68a85 · inbound

CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention cites this paper.

CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T12:54:06.538826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:54:06.538826Z digest=sha256:0d3fff177a87699067f3583b1ec31ac7257be7cefcfdcb7dfafeb24414fd9dee

Observation 3c705765-eeba-40e8-95fc-02ca55f8dee6 · inbound

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks cites this paper.

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:26:59.312672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T01:16:07.252429Z digest=sha256:55e8c05ea907253587fdd87d58795ffd53bff551951736b84b31464fd7321a54