Pith. sign in

Paper Citation Record · LEDGER

Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2402.09063.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.09063 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:32:24.110160Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fc9c4729-bf71-48df-b815-87370bc68e8b · inbound

Fast Proxies for LLM Robustness Evaluation cites this paper.

Fast Proxies for LLM Robustness Evaluation Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T19:32:24.110160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:32:24.110160Z digest=sha256:a5462860d4ad214285aed851bdf9c8e3ad6b2a3fa8aa542c7464ad747875d5c5

Observation cc335f10-b94a-4214-9dc5-b34398eefa78 · inbound

LLM-Safety Evaluations Lack Robustness cites this paper.

LLM-Safety Evaluations Lack Robustness Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:27:21.349571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T01:26:45.402983Z digest=sha256:b83f5003d27d9e26bae3b229a0d84979db45d9126593beccd5533521a6ab26eb

Observation e8e073b0-f459-4c27-a790-27bc91acf10b · inbound

Fake Friends and Sponsored Ads: The Risks of Advertising in Conversational Search cites this paper.

Fake Friends and Sponsored Ads: The Risks of Advertising in Conversational Search Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:24.716762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:24.716762Z digest=sha256:069494e819e8378d99b751d2b778a0cbe6bb13d397321d98210d8019f40e1829

Observation d96455b2-84ed-4436-8021-783c7711c8d1 · inbound

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks cites this paper.

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:50:25.052460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:50:25.052460Z digest=sha256:8c06674a5c42d6725254fbc7079202e9a32774b44c80c259d9e6651e9fb655b2

Observation ba850a58-8388-4051-a10f-7b95ce1f8678 · inbound

Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs cites this paper.

Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:53.129178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:53.129178Z digest=sha256:6474bb2e91700a36be82e2224102b2c7f8af86d76a97f808f821b853403bb889

Observation c0fae038-732b-47ca-9473-9237fc0effa6 · inbound

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation cites this paper.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.718850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.718850Z digest=sha256:784b7c18eda2616094e400280953edf9dd906b0f9c721701b7cbb0bb0b20e0f6

Observation 25011324-4e75-4273-8cdd-0c3901a57082 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:37.155433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:37.155433Z digest=sha256:84f5f8370e0c4aa60e69d24dac90991ff03607d718683ea4f503206482cce039

Observation f863855c-2927-461e-94ff-d7ea1978ab0e · inbound

Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs cites this paper.

Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:21:00.105707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T16:41:52.440793Z digest=sha256:b62abb4f412d274eef65fd9e6652aa4e5768a135a437c689f4c39d4c3b876d92

Observation aeb00d55-7c17-4338-9bac-3b1bd8a0f3d2 · inbound

Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data cites this paper.

Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:21.589235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T05:56:38.042978Z digest=sha256:eff998d94b04039e8871ee3e593adbce7ba19386757b2a98415b20afedd0c10b

Observation d3268f6a-4563-4f02-bc30-06e7152dc068 · inbound

Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes cites this paper.

Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T15:40:36.714300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:40:36.714300Z digest=sha256:fe3b52956db0efc88b04523e1b204ca558adacdf88629acf2a287ed52632187b