Pith. sign in

Paper Citation Record · LEDGER

Sparse Autoencoders are Capable LLM Jailbreak Mitigators

As of 8 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 1 inbound Pith citation observation for arXiv:2602.12418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.12418 v2

Coverage vector

measured 9 of 9 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T23:52:40.912038Z

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-30T20:47:40.813521Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

9 of 9 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved8
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 582e6c86-985f-4869-aa50-b70660a12d98 · outbound

This paper cites an unresolved cited work.

Sparse Autoencoders are Capable LLM Jailbreak Mitigators Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T23:52:40.846642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:52:40.846642Z digest=sha256:65f080e3655cd4b202f9eba6461e3d8bc29d9ce4f58e99212167964725c896ee

Observation acfae772-0729-4c05-a39e-eeaf0f77960c · outbound

This paper cites rights around free speech and freedom of assembly.

Sparse Autoencoders are Capable LLM Jailbreak Mitigators rights around free speech and freedom of assembly

Reference 2

Resolution
malformed identifier
no resolver link, observed 2026-08-02T23:52:40.912038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:52:40.912038Z digest=sha256:c3cceb1c3c8df394f9f382b4051732df7c357be7a486e0b98ef108704e3a5a5a

Observation 1aaf292f-8844-46e5-8042-bc758f4b19ea · outbound

This paper cites Intriguing properties of neural networks.

Sparse Autoencoders are Capable LLM Jailbreak Mitigators Intriguing properties of neural networks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T23:52:40.515937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:52:40.515937Z digest=sha256:d6141e14b762ce1a0f1a380f2b8123d45ba3836c56b5b23f20b08a138bd58332

Observation 7df0063e-b2b4-458e-96d4-16550bff531e · outbound

This paper cites an unresolved cited work.

Sparse Autoencoders are Capable LLM Jailbreak Mitigators Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T23:52:40.697575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:52:40.697575Z digest=sha256:476b71d4b1d7375341e1f87810d0f8d2caa4643154386a5d680f3ff1a786beba

Observation 27cefc0c-5eb9-40a4-90cb-8cacecfa42e3 · outbound

This paper cites — "You should be a responsible Language Model and should not generate harmful or misleading content! Please answer the following user query in a responsible way. {}.

Sparse Autoencoders are Capable LLM Jailbreak Mitigators — "You should be a responsible Language Model and should not generate harmful or misleading content! Please answer the following user query in a responsible way. {}

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T23:52:40.790204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:52:40.790204Z digest=sha256:d9edd771b61e50bcf02d05c4b74aa890b70ace06515e3590b2fc270c65ea0a74

Observation 191d38ee-d039-4b79-aec7-4800c10e19da · outbound

This paper cites AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks.

Sparse Autoencoders are Capable LLM Jailbreak Mitigators AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-02T23:52:40.608795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:52:40.608795Z digest=sha256:1eb089f2d989a52f1e7dd5fcc62a141f47ad894d2cf9b52d5fab9d585c617945

Observation 8c76d597-7b76-4d26-ae18-5bad78f5d46c · outbound

This paper cites RobustBench: a standardized adversarial robustness benchmark.

Sparse Autoencoders are Capable LLM Jailbreak Mitigators RobustBench: a standardized adversarial robustness benchmark

Reference 844

Resolution
unresolved
no resolver link, observed 2026-08-02T23:52:40.329981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:52:40.329981Z digest=sha256:0281d9023b3cf189a293204cb50db96b19dc2a651ba50e2383429b8dbb9ac897

Observation 51025924-d1e2-4e81-bb09-9efbfae16cef · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Sparse Autoencoders are Capable LLM Jailbreak Mitigators Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T23:52:40.420881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:52:40.420881Z digest=sha256:88400aa887f87eb1735dfa69f78f7ff9982dc59ce66689b67999f0a029569263

Observation 1a677b74-7079-44f6-a909-3ad05bc53d35 · outbound

This paper cites Steering Large Language Model Activations in Sparse Spaces.

Sparse Autoencoders are Capable LLM Jailbreak Mitigators Steering Large Language Model Activations in Sparse Spaces

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T23:52:40.237099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:52:40.237099Z digest=sha256:80bd8dc38eb98afddd6e5aef1e852d69d51a659048f87062e4a41ce96ff1bb1d

Pith citing papers

Observation f51b1d39-94d5-439d-874f-899be739798d · inbound

Do LLMs Know Their Vulnerable Scenarios? cites this paper.

Do LLMs Know Their Vulnerable Scenarios? Sparse Autoencoders are Capable LLM Jailbreak Mitigators

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-30T20:47:40.813521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:47:40.813521Z digest=sha256:38561e2e35bbb171edc01392c07eb569e8695402726cf8b685e0a6808abdf617