Pith. sign in

Paper Citation Record · LEDGER

How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2406.05644.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.05644 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:20:23.515973Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T07:46:56.508546Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 800db129-f8ac-4012-9812-1e8d11a6ae4b · inbound

JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit cites this paper.

JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T19:00:24.118281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:00:24.118281Z digest=sha256:e47756e739f9cf865da6d251ff640ceaa62d897fbe8e2eb8cf03a15d7b5279a1

Observation b45fd81d-2375-48f8-b76f-54b70f32a4c4 · inbound

Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models cites this paper.

Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:13.121189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:13.121189Z digest=sha256:9caa404405849e158f31ff50020a6c7b9ad6c36b26e16229d001b48f3695b4b8

Observation 18471b95-f458-4144-99e5-ae2ab6e31075 · inbound

LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models cites this paper.

LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T23:40:38.960198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:40:38.960198Z digest=sha256:5b98dc38fb8626625358c7efb0e95515fc9f616ee50cc0be6048beb78c446723

Observation 7faf2a84-6290-4cb6-8108-dec1e0cbd28d · inbound

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation cites this paper.

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:30.728269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:30.728269Z digest=sha256:522b5aa7bfde8aea823d613e5980f646a6e74b9c3642495bf479c6fdf142256e

Observation 351cdc1d-19bd-4ba0-8c42-0184ba06cfb5 · inbound

AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM Security cites this paper.

AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM Security How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:20:23.515973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:20:23.515973Z digest=sha256:977c7f73899aa25b2e3fa6d5de9ab71ddd8761ad6280cdbb46d51b854dcdd4c9

Observation 2cf4d2f6-e9d5-4033-89f9-35cef040d227 · inbound

Latte: Transfering LLMs` Latent-level Knowledge for Few-shot Tabular Learning cites this paper.

Latte: Transfering LLMs` Latent-level Knowledge for Few-shot Tabular Learning How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:58.246436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:58.246436Z digest=sha256:305536480ac5750a0dc410d0b0bc1ee4cd3f3cbb1378207841bcb0da6afd2b37

Observation 0b270832-d925-4afe-950e-90d540832e67 · inbound

Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement cites this paper.

Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.350773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:45:55.350773Z digest=sha256:bc7d410ac5cc363a4efd4920d01b9c82236d68aeb244c7f5e850cbbda291788f

Observation 7870e03f-04b4-4169-a6d9-619b08e956b8 · inbound

GeneBreaker: Jailbreak Attacks against DNA Language Models with Pathogenicity Guidance cites this paper.

GeneBreaker: Jailbreak Attacks against DNA Language Models with Pathogenicity Guidance How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:24.614398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:24.614398Z digest=sha256:046e9108668d0f410a0a5de1355d7e8c764e1c9a21f4de3c88b26e7c048f7a56

Observation 429d95b4-f72c-454b-a70f-8be9adf9e52c · inbound

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models cites this paper.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:33.346536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:33.346536Z digest=sha256:ea428b9c8705f29b0e0dada463cefc07dac78696cba4020cf5c8afa358dfb8a0

Observation 51d00b66-1e49-4927-b73b-35dfe637e7e1 · inbound

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers cites this paper.

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T16:28:53.895617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:28:53.895617Z digest=sha256:517225a99010525854a384442cd8af9c1123169f96c0e3f404cd4ce4b4973c31

Observation a6e098f9-c354-479b-b8f2-9964ef0d3085 · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.500858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.500858Z digest=sha256:a21086276dfb5a1470c5843d68a6a6ff4fb434fcbc7c014a0fa3133bfd20fb19

Observation 7714dea5-9a52-4358-9262-6a615eea9535 · inbound

Why Do Large Language Models Generate Harmful Content? cites this paper.

Why Do Large Language Models Generate Harmful Content? How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:21:04.470748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T15:31:13.545599Z digest=sha256:9f05283afe790e30ac22bab5b9ef6a9b58573e74a272e5693e5101bbb4cb4020

Observation 3eaf97ba-80cc-4432-8313-40a0fbf8663d · inbound

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models cites this paper.

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:49.276055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T18:43:12.298529Z digest=sha256:cf3fccedcdd60e40cae176657e7e46b6f701538d414be85b99ba3bd67d30e662

Observation 6df54325-e926-4bbe-8a01-62e1217e5e4d · inbound

Before the Last Token: Diagnosing Final-Token Safety Probe Failures cites this paper.

Before the Last Token: Diagnosing Final-Token Safety Probe Failures How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:38:00.112172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T21:37:02.439454Z digest=sha256:4e69e8223e25e5fc2103d323895a5e6d0194b35354e56dd246324b0e9a73be1e

Observation 38a18b8f-01ee-495f-956b-76c41180bbf5 · inbound

SafeSpec: Fast and Safe LLM via Dynamic Reflective Sampling cites this paper.

SafeSpec: Fast and Safe LLM via Dynamic Reflective Sampling How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:59:33.720230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T17:23:06.382761Z digest=sha256:0cdc7fe18546518a34d52329011164b1098f26380ec3477feda03febc7b4e7ee

Observation 3fbb7da7-e5fb-447a-8874-2ff1e1274ff6 · inbound

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice cites this paper.

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T07:46:56.509857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T07:41:35.741974Z digest=sha256:532729cf859a0ebc3ebde5453f527e2ff58e3398a82e9efc4c9762ac02371b48

Observation be95769d-5559-4498-b31b-3c400c93723e · inbound

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice cites this paper.

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:06.547479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:06.547479Z digest=sha256:cdf9f431656d346118e86b50e308f745f02a95e42ad3211dde2ecdb2c00ce3ca

Observation bfd23706-d0de-4641-a148-94c13ce1ec17 · inbound

Verbalizable Representations Form a Global Workspace in Language Models cites this paper.

Verbalizable Representations Form a Global Workspace in Language Models How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 186

Resolution
unresolved
no resolver link, observed 2026-08-01T23:15:30.920135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:15:30.920135Z digest=sha256:0325a0a87d705bb966fde15cb3bcde7ae7ecc0daa02ef1ac4f38b888956e69b0