Pith. sign in

Paper Citation Record · LEDGER

How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2406.05644.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.05644 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:20:23.515973Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T07:46:56.508546Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 800db129-f8ac-4012-9812-1e8d11a6ae4b · inbound

JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit cites this paper.

JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T19:00:24.118281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:00:24.118281Z digest=sha256:b26549f2236f97283d8c84a796e93f9eb8aabc63fd3254e4046a00c0f930c051

Observation b45fd81d-2375-48f8-b76f-54b70f32a4c4 · inbound

Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models cites this paper.

Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:13.121189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:13.121189Z digest=sha256:824e52b6521d36d1bfc365b355deb0913b730c63ed6b5c4916a5410e8b75b713

Observation 18471b95-f458-4144-99e5-ae2ab6e31075 · inbound

LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models cites this paper.

LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T23:40:38.960198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:40:38.960198Z digest=sha256:e341e32ce9bfe6434810221b1ed8c212b1ac5cddcdf5a0f7b22e33aebce05874

Observation 7faf2a84-6290-4cb6-8108-dec1e0cbd28d · inbound

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation cites this paper.

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:30.728269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:30.728269Z digest=sha256:3bce285214e18a124e0dc7ccc234e786682e1dc80d294252314d7270381b8d33

Observation 351cdc1d-19bd-4ba0-8c42-0184ba06cfb5 · inbound

AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM Security cites this paper.

AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM Security How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:20:23.515973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:20:23.515973Z digest=sha256:94269757269a4e41ab76767fcf2efcf3b873c11371d2afeded7001a7a2ee5564

Observation 2cf4d2f6-e9d5-4033-89f9-35cef040d227 · inbound

Latte: Transfering LLMs` Latent-level Knowledge for Few-shot Tabular Learning cites this paper.

Latte: Transfering LLMs` Latent-level Knowledge for Few-shot Tabular Learning How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:58.246436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:58.246436Z digest=sha256:d00acfcf9163ed11510599b870cca22dd3a2fbb97a74332f05ab75d205c42634

Observation 0b270832-d925-4afe-950e-90d540832e67 · inbound

Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement cites this paper.

Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.350773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:45:55.350773Z digest=sha256:db688604637dd1a373d6bfab100005026bdaf3ad938ffe6130bb481980910716

Observation 7870e03f-04b4-4169-a6d9-619b08e956b8 · inbound

GeneBreaker: Jailbreak Attacks against DNA Language Models with Pathogenicity Guidance cites this paper.

GeneBreaker: Jailbreak Attacks against DNA Language Models with Pathogenicity Guidance How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:24.614398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:24.614398Z digest=sha256:11582ca5ee478c3c31b75a629514ffb67f82ff9aea5119e87e9648b7f73909d6

Observation 429d95b4-f72c-454b-a70f-8be9adf9e52c · inbound

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models cites this paper.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:33.346536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:33.346536Z digest=sha256:679c1d9f8650902682a9916230920c506b381930a91b9f76abd092ebe4155d09

Observation 51d00b66-1e49-4927-b73b-35dfe637e7e1 · inbound

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers cites this paper.

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T16:28:53.895617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:28:53.895617Z digest=sha256:3557a0680cbdab0f41541706e0657e0710e1ea4636f71cb159b1fc877cc73f44

Observation a6e098f9-c354-479b-b8f2-9964ef0d3085 · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.500858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.500858Z digest=sha256:4e69d2721a09561d6a5cf07cd9eef760e40c809bbcda3395529cbf485fdc61d7

Observation 7714dea5-9a52-4358-9262-6a615eea9535 · inbound

Why Do Large Language Models Generate Harmful Content? cites this paper.

Why Do Large Language Models Generate Harmful Content? How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:21:04.470748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:31:13.545599Z digest=sha256:1e0f62508bb2f280d8e2060a91759bd8b99f7f4efd6a474b92a67594d11cc6fe

Observation 3eaf97ba-80cc-4432-8313-40a0fbf8663d · inbound

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models cites this paper.

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:49.276055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T18:43:12.298529Z digest=sha256:e193979065697ca409064c7b25fab2bf36e65474ebfaead41a67b505f5caaf37

Observation 6df54325-e926-4bbe-8a01-62e1217e5e4d · inbound

Before the Last Token: Diagnosing Final-Token Safety Probe Failures cites this paper.

Before the Last Token: Diagnosing Final-Token Safety Probe Failures How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:38:00.112172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T21:37:02.439454Z digest=sha256:d6a3b7d5297cffb4b942be225060f2dbf7b9c189ffaea985f1d9c251725fd4de

Observation 38a18b8f-01ee-495f-956b-76c41180bbf5 · inbound

SafeSpec: Fast and Safe LLM via Dynamic Reflective Sampling cites this paper.

SafeSpec: Fast and Safe LLM via Dynamic Reflective Sampling How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:59:33.720230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T17:23:06.382761Z digest=sha256:5b607cc872a00dceea42a5a77fbe63f3f63cd1d3fc2581e254ad5eb4cafabd09

Observation 3fbb7da7-e5fb-447a-8874-2ff1e1274ff6 · inbound

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice cites this paper.

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T07:46:56.509857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T07:41:35.741974Z digest=sha256:f5e0055fc31b92ffc712f4f7757fe1dffa3a911e3f871a3c2270d177252435f1

Observation be95769d-5559-4498-b31b-3c400c93723e · inbound

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice cites this paper.

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:06.547479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:06.547479Z digest=sha256:932a9707085f72484829d1fd01394948316f8d925ddd843cb75e66122cfcce51

Observation bfd23706-d0de-4641-a148-94c13ce1ec17 · inbound

Verbalizable Representations Form a Global Workspace in Language Models cites this paper.

Verbalizable Representations Form a Global Workspace in Language Models How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 186

Resolution
unresolved
no resolver link, observed 2026-08-01T23:15:30.920135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:15:30.920135Z digest=sha256:2b9dda995ac00c023f7b07e5eef7066a4b4b7535eed1607c5832f22d09578261