Pith. sign in

Paper Citation Record · LEDGER

Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2405.18166.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.18166 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:55:13.112842Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T11:53:03.533122Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b2444566-31da-4d83-80e9-e5ad8712331d · inbound

Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models cites this paper.

Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:13.112842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:13.112842Z digest=sha256:3e8bea5436e1ad3f46a90921c27c4f9231eb3a9012ff2ac3fb46477223f5db56

Observation ea3b3048-abec-45b6-842e-8574a3f63a75 · inbound

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense cites this paper.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.181315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.181315Z digest=sha256:104d4e2f81dc4d1f2542771375375e68d93304d08bfde3a18c17db16fc4db25f

Observation dc5a12ec-80b6-4460-b6ec-113c69356cf7 · inbound

On Almost Surely Safe Alignment of Large Language Models at Inference-Time cites this paper.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.989228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.989228Z digest=sha256:e8db0d8b71a3ce85a2229bfcd20179998fbff3165f597ef907fb4baf9fac329e

Observation 96d8eb54-9b08-4fa5-a3d1-5330150667ef · inbound

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation cites this paper.

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:30.717668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:30.717668Z digest=sha256:b25167dad683c2151ff76364629f7f61c0950ebcec5ea6857dd42b371cdd4b66

Observation 968a4e6f-0bb3-495f-b052-a572df396cbc · inbound

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors cites this paper.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:33.763079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:33.763079Z digest=sha256:38900c25fc9b9c8858f704156d6acd7ae849fdf6a4f9d6ff2eb32831b4394f48

Observation 011b081c-b633-4b13-b08b-47f08cc71277 · inbound

GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace cites this paper.

GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:34.599908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:34.599908Z digest=sha256:644c83e55202a57e8d7c75f69818cc173c090f9c7022b04fdd35ad0337f2c0ed

Observation 3e9739da-bfe8-407d-8fbb-0b0a745ba207 · inbound

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment cites this paper.

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:53:03.534832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T11:52:36.688263Z digest=sha256:63a58f2cefd6ca56843dcb100d62db1bcce37912429c369e69d773949bd7d8e0

Observation 3b6d11f3-dffd-46e9-91d9-fd8f4ce197b0 · inbound

PUZZLED: Jailbreaking LLMs through Word-Based Puzzles cites this paper.

PUZZLED: Jailbreaking LLMs through Word-Based Puzzles Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T05:43:39.775290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:43:39.775290Z digest=sha256:58bf36c02b50422dba58965e7a0c3e6f535d7a43504f2af2a347a59d927f4709

Observation 4e9e3dab-f6bc-4b05-83ed-17188a1e8e3f · inbound

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring cites this paper.

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:31:19.440660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T22:28:41.134253Z digest=sha256:9690e112a7b833cb0ce3623b55b6076826806b98ea75157299a25b362727d43c

Observation 7d3e94ee-1bdb-4344-8ae2-4535f55d15db · inbound

Silencing the Guardrails: Inference-Time Jailbreaking via Dynamic Contextual Representation Ablation cites this paper.

Silencing the Guardrails: Inference-Time Jailbreaking via Dynamic Contextual Representation Ablation Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:05:58.322946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T17:48:09.607231Z digest=sha256:9b7db56f7bbbed3d66611f9cc54994b2cd58a54ad95bebd15a14febaaa21f16b

Observation 106f1c68-28ca-4dcc-9156-33c3e458f114 · inbound

Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types cites this paper.

Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:31:00.365870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T17:08:25.471462Z digest=sha256:3186a08e5fd539c0b32b739c701462991eabf4ae6e687df98f295fb18de326ed

Observation 7645e6d1-bd18-4654-b0f6-c8e614faa764 · inbound

LLM Safety From Within: Detecting Harmful Content with Internal Representations cites this paper.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.143551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:34f0f0edd77fe12cfc56c1a474f6cb95a5d34414f1034acdea3df4420961ac7f