Pith. sign in

Paper Citation Record · LEDGER

Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2402.09063.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.09063 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:59:24.716762Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cc335f10-b94a-4214-9dc5-b34398eefa78 · inbound

LLM-Safety Evaluations Lack Robustness cites this paper.

LLM-Safety Evaluations Lack Robustness Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:27:21.349571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T01:26:45.402983Z digest=sha256:c55df6fcfb1d6b0ab48cdd45e1f70eccdcbb96bec12f4add5f03d63d792fbae3

Observation e8e073b0-f459-4c27-a790-27bc91acf10b · inbound

Fake Friends and Sponsored Ads: The Risks of Advertising in Conversational Search cites this paper.

Fake Friends and Sponsored Ads: The Risks of Advertising in Conversational Search Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:24.716762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:24.716762Z digest=sha256:069494e819e8378d99b751d2b778a0cbe6bb13d397321d98210d8019f40e1829

Observation d96455b2-84ed-4436-8021-783c7711c8d1 · inbound

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks cites this paper.

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:50:25.052460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:50:25.052460Z digest=sha256:842d64c1d7fc00c333e65f8ce24d1ff30c67d957a04ecea2e0b6f5b0eec6f0c2

Observation ba850a58-8388-4051-a10f-7b95ce1f8678 · inbound

Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs cites this paper.

Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:53.129178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:53.129178Z digest=sha256:714d46bc707deabf5b09514cb7d8bc44429b6c8ba8baaffb150c8a96f8d752fe

Observation c0fae038-732b-47ca-9473-9237fc0effa6 · inbound

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation cites this paper.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.718850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.718850Z digest=sha256:784b7c18eda2616094e400280953edf9dd906b0f9c721701b7cbb0bb0b20e0f6

Observation 25011324-4e75-4273-8cdd-0c3901a57082 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:37.155433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:37.155433Z digest=sha256:d9ae503878414f746e28751a57360fc424f1fea5111fd7d71a97650a20965c8c

Observation f863855c-2927-461e-94ff-d7ea1978ab0e · inbound

Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs cites this paper.

Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:21:00.105707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:41:52.440793Z digest=sha256:bfb1f273937f3f9805ca0e97b6c6b9494f77c1025dfd25c1cb7651f4d79f60a7

Observation aeb00d55-7c17-4338-9bac-3b1bd8a0f3d2 · inbound

Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data cites this paper.

Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:21.589235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T05:56:38.042978Z digest=sha256:940daef8ece41e508d0884d3d64ad661bf9c4c2845b1a3dfb7b5b172e4b4c8c1

Observation d3268f6a-4563-4f02-bc30-06e7152dc068 · inbound

Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes cites this paper.

Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T15:40:36.714300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:40:36.714300Z digest=sha256:61a861d6c46187ffb4839de9859ffba480fac76840c1955d46615d3710f78258