Pith. sign in

Paper Citation Record · LEDGER

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak

As of 15 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 2 inbound Pith citation observations for arXiv:2412.17522.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17522 v2

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:31:31.746425Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:13:58.260498Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 05eed926-717a-4a11-9f32-31b84cfff728 · outbound

This paper cites an unresolved cited work.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:31:32.367338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T05:31:31.547505Z digest=sha256:ddf89b0069d9a1b37fd3aed99e9abe7549dd7c73e106828db8ffc9fc3999621d

Observation d86034da-2d3c-4d3a-86d2-3a939443e31f · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.552515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.552515Z digest=sha256:1727dfb6d9d9f29c4e4714a241bef86b3d4e946f668f5a393426e89faeaa88ae

Observation 78d196db-ad97-49f7-871d-c73f7150d652 · outbound

This paper cites an unresolved cited work.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.557406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.557406Z digest=sha256:2eaf33a66adf55ca3f71022803f7f7ccb25a43ce4aec63c767b18a2f0eadc942

Observation 258167ae-6509-46a5-a8b7-a21c722cad5d · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.562431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.562431Z digest=sha256:a884f47bd37d3097d1b712fae1535f5bffab0be1215f5104a6bf2313f9943aa7

Observation dbdd6212-9106-47a2-85fd-dbd762144e02 · outbound

This paper cites Plug and Play Language Models: A Simple Approach to Controlled Text Generation.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Plug and Play Language Models: A Simple Approach to Controlled Text Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.567234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.567234Z digest=sha256:f4273e3b5917be91efd8c3ca2664e4ca1fe70399f463b54e8a4f5b723ade996f

Observation e8d74daa-42c4-4fdf-a5d3-e7f0ccd4df6b · outbound

This paper cites The Llama 3 Herd of Models.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.571939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.571939Z digest=sha256:8160cd355b73f2eddcadd2141ae0b81f60c6dbf3f4a91f648a386f586c7d32d3

Observation 66422da4-9284-4434-adf0-a283237a2e44 · outbound

This paper cites Attacking Large Language Models with Projected Gradient Descent.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Attacking Large Language Models with Projected Gradient Descent

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.576750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.576750Z digest=sha256:0e5e6b537735adf50366b881b0ef09a12d8455a9d18e3ecc6f23086d82e79f89

Observation 3bf966e1-90ab-4a5c-a92e-c23584959e55 · outbound

This paper cites DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.581240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.581240Z digest=sha256:949c9436a565a62834c99ceea41fcb7c1b3a46d8fc5b7a50c919144a3391feec

Observation a625f936-e936-46ac-aa9b-eb391abd8c0d · outbound

This paper cites Gradient-based Adversarial Attacks against Text Transformers.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Gradient-based Adversarial Attacks against Text Transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.585911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.585911Z digest=sha256:a6509a76f36ed6fe851d40439a217d4dbe22fe366a7f586ce9d5c2c3e1cb124d

Observation da67d322-0dc0-48f2-9877-299b2f4e60d8 · outbound

This paper cites COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.589895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.589895Z digest=sha256:51d9ece1abc22ccc7bf0b6ff8fbc0703040b6f673918c2070432af832fd25e71

Observation 87a80c1c-6819-4178-863c-700e4be298ea · outbound

This paper cites an unresolved cited work.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.593737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.593737Z digest=sha256:c75d3518e0262b3e9c71b77a08b3a8f38ec114df71b19e726c8738dd311b8ae7

Observation e3bd28ae-2b37-46ff-85b3-1c348a424fd6 · outbound

This paper cites DiffusionBERT: Improving Generative Masked Language Models with Diffusion Models.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak DiffusionBERT: Improving Generative Masked Language Models with Diffusion Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.597821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.597821Z digest=sha256:7ed7a30a27c6159640d7e3c45865a28dfe90f0a399d1d5c9afeb77a33909c3ca

Observation cc28aa36-56b9-442a-a3b2-78a3b6ede325 · outbound

This paper cites an unresolved cited work.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.601771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.601771Z digest=sha256:3ee1ff77df91f6b61e74093c050491f757463c3b5e971892b0a37e1d95044a8c

Observation 85f2a65b-4279-44a0-961b-73e8828e3f89 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.605922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.605922Z digest=sha256:b7ff7926b071dd00bbaafc72ba067f04f27a3742160e74c3343224c9cfb2dc6c

Observation 35f622f5-f28a-4d20-8c9b-2b2922760e8a · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Categorical Reparameterization with Gumbel-Softmax

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.610328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.610328Z digest=sha256:77107e207a88add7c8de30283ef22182e43c6ebba0aa6de0d792ae640784809a

Observation 0904a54e-2896-4d2d-9875-522f8232332b · outbound

This paper cites Mistral 7B.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Mistral 7B

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.614678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.614678Z digest=sha256:441837d175a724ac82956390e5a3be55f39c5a6b18d38ee943cf890538f87105

Observation 68185f2f-8481-4f23-bf10-81b2a288b926 · outbound

This paper cites an unresolved cited work.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:31:32.327797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T05:31:31.618979Z digest=sha256:cead165c46e9a6b49f146b02de99eb90ae0f9cbda8f01c94d26cbb3fea4fa042

Observation 9142b51b-2d74-46a8-87c2-d5258cd8c6a7 · outbound

This paper cites an unresolved cited work.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:31:32.313078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T05:31:31.623159Z digest=sha256:ed4a4eded2f4ce405866b78ba5c4bedbbc337894637b03afc4ed01c7c9d45d22

Observation b98a1bfc-e124-4bc0-ab30-ffa5a97e37a9 · outbound

This paper cites an unresolved cited work.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:31:32.297780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T05:31:31.627132Z digest=sha256:9384166850bad3dbac88a31fa5e180f7cdb032f4c2cf2daa948db17843f1583a

Observation 77ae4c2d-05d4-4ae1-aec4-21d7c24972ae · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.631402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.631402Z digest=sha256:4b96be6eb2f2ffc56ec4ded8f4304bcaf7ca45a0a7d970457f510898c67a5ff4

Observation 53768670-84d1-4730-be5b-172b034e4a1b · outbound

This paper cites Discrete diffusion modeling by estimating the ratios of the data distribution.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Discrete diffusion modeling by estimating the ratios of the data distribution

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.635839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.635839Z digest=sha256:bace8f3eeb669cd8cb9ea1dd5dc00be6b43a99f15961af8b9aaca5bc83e900bc

Observation 6d435254-a92a-4b2e-bf17-e1230535c943 · outbound

This paper cites an unresolved cited work.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:31:32.273069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T05:31:31.640070Z digest=sha256:1a9a60d6b06477071f71d6f084daa650d77fdcfca7a33f8094a4876a47408934

Observation ea4e34ad-780a-441c-b28b-86fd8c9742fc · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.644331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.644331Z digest=sha256:e7697bc0001cda7e9565c7a75ff9370199fe3415df22547aaeae603b873f0f05

Observation f72979ca-8f07-469a-a2e3-dfbd03e41ae3 · outbound

This paper cites an unresolved cited work.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.649193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.649193Z digest=sha256:6ed061a221c23d3dd7045471847cf2610fcf5b5b41be2da13fa1e9c711ddab85

Observation 70337389-7629-4802-b532-15ef011f6c38 · outbound

This paper cites AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.653316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.653316Z digest=sha256:4fdfd33d0f9d6e54cf631b523fa3c4129130c05592ad259651fed15a2affb63c

Observation 6820cef2-53ea-4c85-bf2e-f59dea371f3f · outbound

This paper cites Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.657432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.657432Z digest=sha256:3f1e5ca780db1364e68c3ca712c771acddbfafb7415d588b3781c7765cc8781d

Observation 5c9f58f2-c56e-4cdf-873c-0ceb2da7249f · outbound

This paper cites AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.661915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.661915Z digest=sha256:0d24e9f24ddfcbe130f8105d7fa694fe03d5c32a8ba82e1719921798f52b9992

Observation 90fed09a-a740-4dd2-aea9-1d53338de80e · outbound

This paper cites an unresolved cited work.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.666471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.666471Z digest=sha256:15523c76bdcbece024c9d6710e2891c0fe0d1f59f07bf8a106358be47eb758d4

Observation 0dce9e06-85cc-4585-868c-2b24b4433b98 · outbound

This paper cites an unresolved cited work.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.670815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.670815Z digest=sha256:803f62ce5fd7765e7e1f3d09a0995197eb6fc7d652b7cc5c156aaf4a3f3aedcd

Observation 6013bc43-a3d9-40b3-b904-79138c4dfd7a · outbound

This paper cites an unresolved cited work.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:31:32.232860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T05:31:31.674786Z digest=sha256:ab3204d2f1426df7375c4d27c0d28cc005e545f07f696953d8c48b136b40b072

Observation ac8622cd-3341-4e7c-8dc7-02b0d9625542 · outbound

This paper cites Harnessing the Plug-and-Play Controller by Prompting.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Harnessing the Plug-and-Play Controller by Prompting

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.678401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.678401Z digest=sha256:e3b8e794ce71074bc3ebf798f0aae3bb8d77c19d86f6674b4640c8d978c4f12d

Observation fba8892d-4e71-4bd4-8f40-9fa3d6560746 · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Aligning Large Language Models with Human: A Survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.682237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.682237Z digest=sha256:678a29f659d5b598150a386cac0d683f9535ec26e700a3f65adee11dd8a1dccd

Observation cd4f495d-a951-47e0-9bb6-39d5b07c78c7 · outbound

This paper cites an unresolved cited work.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.686323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.686323Z digest=sha256:efd4e546d1c31bae6c598f323deae6e2e15d61c8ecd4d643bf8f227c21626ff5

Observation 2554d994-edb1-44ee-bc37-68b81b04d58f · outbound

This paper cites Gradient-Based Language Model Red Teaming.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Gradient-Based Language Model Red Teaming

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.689991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.689991Z digest=sha256:3131ca4508cd199342efce5ebe286f70fbb7b91149ba7b08cb668e6bc9b81bae

Observation 0d762d3c-4214-4523-b47c-f83939a6c613 · outbound

This paper cites an unresolved cited work.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:31:32.209076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T05:31:31.693751Z digest=sha256:d11fedd26a765a3d2a403c7dbe8a3218e06f83722d795805bdce57facc547834

Observation b92edd19-5e98-4337-91f6-107f28fbfbc0 · outbound

This paper cites Jailbreaking as a Reward Misspecification Problem.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Jailbreaking as a Reward Misspecification Problem

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.697628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.697628Z digest=sha256:da338f063b7868d5898adbfe4b66ec0986c64547300569f27ddf33c5d095e9ca

Observation 002b8ce0-ef96-4d78-80de-b2de3873f9af · outbound

This paper cites an unresolved cited work.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:31:32.192862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T05:31:31.701971Z digest=sha256:ffc747d49cdb17a30798bc08e1936fff62610c638076b70bb2704c37eda0ae8e

Observation a5b11e94-ed15-4d34-abee-79438d50a299 · outbound

This paper cites DINOISER: Diffused Conditional Sequence Learning by Manipulating Noises.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak DINOISER: Diffused Conditional Sequence Learning by Manipulating Noises

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.706019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.706019Z digest=sha256:6f2dfc592d2458a09a39174f59884b471d6b80da485e917619f2f0342b45dad8

Observation 51edbe05-782d-44f1-9cf5-0b0c48e7f1b8 · outbound

This paper cites GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.710393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.710393Z digest=sha256:1161ad5790669ca64fa8908a7780b920a8b33b27b673429e2caefb3c8ff2670e

Observation dc37adfd-afe1-43ec-b639-b5a76c831ed9 · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.714580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.714580Z digest=sha256:77f589d0b2df32d281b5c49146a2adf79306072f44d010fd4b10201866bd59ad

Observation 5ce9537d-bafe-470d-a4f3-6956a74756c9 · outbound

This paper cites PAWS: Paraphrase Adversaries from Word Scrambling.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak PAWS: Paraphrase Adversaries from Word Scrambling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.718925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.718925Z digest=sha256:d9fe4f75f3d9b51b354c69c90f689da3fd194b71f4d1873d85f3783549f55b12

Observation 792fd9b5-499f-4abc-946a-4f9c4d7886a5 · outbound

This paper cites On Prompt-Driven Safeguarding for Large Language Models.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak On Prompt-Driven Safeguarding for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.723417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.723417Z digest=sha256:6592aaa93add57366f1c69ac04616bd4995e39c87ecedec93d9466e4d768ff3f

Observation 1d167eb0-6339-4370-b193-4b6957f5d2b3 · outbound

This paper cites AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.727827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.727827Z digest=sha256:cfe6b5fba2524c0d4f2e45774ae95928d42db123295be2495b20a17bb45076c3

Observation 4d06e451-8927-46e8-8431-35398126ba2c · outbound

This paper cites an unresolved cited work.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.732594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.732594Z digest=sha256:09515aab247babb68af3b5ca57d85a5c3195203dcdc507023f6a14ea324c4ada

Observation 1c5f40b0-75db-4879-8887-5808a8f13fd3 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.736732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.736732Z digest=sha256:d764eab06400fa0b9f18762ade9f26ee6f2ce760d69e2b7df475f72c8dcf569a

Observation 4edaf3b3-6454-45a2-bb4d-aa914f95184e · outbound

This paper cites online" 'onlinestring :=.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak online" 'onlinestring :=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.741097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.741097Z digest=sha256:bcb89f32155e66c0346301672a8629bb2555f2d8a02c73241d037d44c91d96ac

Observation c12655af-315a-4027-b665-8530a972e3e1 · outbound

This paper cites write newline.

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak write newline

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.746425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:31:31.746425Z digest=sha256:1ad3b415f2a548abe67789d3fdf85aadba634485bd80d9cd0d5de3225cfe7578

Pith citing papers

Observation 53369d06-f6a7-46af-a47e-8343964a5297 · inbound

Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment cites this paper.

Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T18:13:58.260498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:13:58.260498Z digest=sha256:ad055edf1304cdc8d68094e92600889dcd9afcacd73a071352c143d712927cea

Observation 2bce4872-3f11-43f0-a1a1-be8949728b2f · inbound

Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models cites this paper.

Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-06-26T04:38:59.152801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T04:35:51.583460Z digest=sha256:c830a4de8e4a81a2476f02c8da3850cf836a71adeca6bf2898a6483345f47874