Pith. sign in

Paper Citation Record · LEDGER

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning

As of 11 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2501.07959.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07959 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:35:52.915101Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 51987e9b-db86-4e42-85d4-7ed965eb8c5d · outbound

This paper cites Llama 3 model card.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Llama 3 model card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.697345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.697345Z digest=sha256:57e9315f7b4bd6a922808924526babf9105d1393274a174ca9bd996aff40a57b

Observation 0e399931-dc77-41bf-a8d0-8557088251a1 · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Detecting Language Model Attacks with Perplexity

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.701802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.701802Z digest=sha256:b40529619612e72c0d15ab81f06e3f3aadf2623238af8ad1c75cf9cd30ad0391

Observation 3af79303-da0c-41c9-affa-cbf70b6b7e2c · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.705776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.705776Z digest=sha256:d68afe0a6823ce0a9649bfc8f6ed9f8d19a2f754d1327909ceab431633b670a6

Observation 094b98a7-515f-4d18-b11a-c305be6ca13e · outbound

This paper cites Many-shot jailbreaking.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Many-shot jailbreaking

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.573329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:52.710220Z digest=sha256:b23d29625a9f438ad861ed5eee75cac3560a88862b7bc3783a0b17fe4cb27920

Observation 6bb033e1-ce8e-48c9-a83d-809d6234bd9a · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.714194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.714194Z digest=sha256:b0a1e0d6e226ba658a546ea1eb0edd9d1868ccab447de2c42b9f6add3a386bc8

Observation b95c695c-de00-4f27-8d69-e1594b0a73d9 · outbound

This paper cites Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.719204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.719204Z digest=sha256:96b30e79e168b0402259db8328be7f651f64a1e6f7ee99d9e50000aa78427685

Observation bb0f603f-c2e7-4512-bf45-3b8ccb3bf130 · outbound

This paper cites Language models are few-shot learners.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Language models are few-shot learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.724012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.724012Z digest=sha256:b793ad147c33ffbb7b861051fe99e82cc63f4956c616697ff2f7fe1fe9f338e3

Observation 62be13bd-c7b2-43f6-9867-100bd8c422fa · outbound

This paper cites Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.727788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.727788Z digest=sha256:7c0705b36f60799b8c1ca1c0d0527db9f5305a2a0b881fa5957be755cc45118b

Observation ddfa3eb9-4a42-48a3-abe4-935f61abb4e2 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.733088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.733088Z digest=sha256:3435440efc82bfec1e506b44ad416ec1a990ca5272bee7f69e9212d1909854d6

Observation b6413099-72d0-4e21-b4d8-71bf0ae35d1e · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.737657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.737657Z digest=sha256:7959121bd80fa760390be794bf39111dd596456d71305a059ed4161790d623c8

Observation 2138c5fa-c552-429f-a962-5fed6f1913ae · outbound

This paper cites Combating misinformation in the age of llms: Opportunities and challenges.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Combating misinformation in the age of llms: Opportunities and challenges

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.742128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.742128Z digest=sha256:9efd4140a7644f8c48972264b802321528ebba28c9bbf0e71119ba49d146c83b

Observation 3d825ee4-f14e-40cc-a7dd-8430b73b6ff0 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.746178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.746178Z digest=sha256:8755c1507a672a7f1a82fbdc9ce33aabffcb7467abedaa1a1a9714e94aec896c

Observation 5a5e26a7-1a84-457a-9088-1aecc83d8781 · outbound

This paper cites Multilingual Jailbreak Challenges in Large Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Multilingual Jailbreak Challenges in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.750811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.750811Z digest=sha256:ebbd4a9549e48bdaceb5d0eb38b99be2be124d64c71197e6afefb494c8fcf7a0

Observation 7a878be1-1938-426c-8f49-4cc7ed5d7ebb · outbound

This paper cites Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.755719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.755719Z digest=sha256:8fe9c7758d20e3d1c393937c46cb59d57416569f019eea5b9a551144479c0c05

Observation fc414b65-0879-4503-8536-00c59249ef86 · outbound

This paper cites The Llama 3 Herd of Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.760004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.760004Z digest=sha256:26029428a689608d76a6bd29fe56b5693cd380ffdecd452f76cf84e747087636

Observation bc2925ad-506c-4846-a1f8-ed75d3afb945 · outbound

This paper cites Red- teaming for generative ai: Silver bullet or security theater? In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pages 421–437, 2024.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Red- teaming for generative ai: Silver bullet or security theater? In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pages 421–437, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.543563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:52.763569Z digest=sha256:8a1de8033941e957956445bcf56e03496d2cafba382ad6538cea5ef54591610d

Observation 86a9f30c-9aee-4579-b0b6-5d765cd04cf8 · outbound

This paper cites Badllama: cheaply removing safety fine-tuning from llama 2-chat 13b, 2024.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Badllama: cheaply removing safety fine-tuning from llama 2-chat 13b, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.531168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:52.767401Z digest=sha256:23682dbdbb74f622d406b8d2ed477cc21f1370dfe38b4de547eb5fddf6d8f50b

Observation 4aebbcbf-ea7f-4815-80f5-81b755a6aee6 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.770507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.770507Z digest=sha256:4b8b0f6cbc3f16529c89cf686a6007e955c6797922d305d90d819d807c8b1c94

Observation 1a230f98-a27f-41a7-9c84-bd66a97c4009 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.774210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.774210Z digest=sha256:e4b5152bf3d81bc0ffa1d65074d416cec786727d93ab498b8cf3fc57ddbadf8f

Observation 641c1b60-bddc-48c9-aeeb-6d8c3eeda8db · outbound

This paper cites Perplexity—a measure of the difficulty of speech recognition tasks.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Perplexity—a measure of the difficulty of speech recognition tasks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.517940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:52.777767Z digest=sha256:af02b5a68a2dd9e3e38d35a5f5aab6dfe383dc939cbd8ca129a422220b8c6721

Observation b52040c0-11c3-4c80-9989-2315068e87fa · outbound

This paper cites Mistral 7B.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.781447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.781447Z digest=sha256:a7b0066c05e0c8e030eaed85f9714c1994a314758b81a4f436f6f1e1dde719ba

Observation a1af48f6-8b91-426a-b038-02fffdf6b462 · outbound

This paper cites ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.785478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.785478Z digest=sha256:bf2f4a3486250a52db709321b60a325b88c7237e18c1d0959efae280cebaa65a

Observation b5580c03-ad9d-4d82-a648-b82d36d9a824 · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.788923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.788923Z digest=sha256:22ca52eeb9001d52d34137a813a0377f0ba2005258afc6b9423285e5b6573c8d

Observation d31ed7c8-b58b-43ea-af7d-f97174656187 · outbound

This paper cites Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.792439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.792439Z digest=sha256:a5cef697ca9ad72019d16caa9c375f662083dca84ac81b58311e2dd608459bd5

Observation e59b2646-f26c-4178-a01d-b79d06bbc569 · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.796233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.796233Z digest=sha256:aea9d9378ea9e9171e4835f8c83b11fff4f5ebc4e590a240a4cf93141a55d862

Observation 774b9cd4-2bb3-44be-9d77-7e489c0def0f · outbound

This paper cites RAIN: Your Language Models Can Align Themselves without Finetuning.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning RAIN: Your Language Models Can Align Themselves without Finetuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.799789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.799789Z digest=sha256:abf97aa6b74f434aa66d9b768840f9f69ed1578233532fee26db75076daf43aa

Observation 6d3a996f-7d66-422b-affe-4e8ac9491e96 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.803418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.803418Z digest=sha256:b8d2a067aff3f4d860a79d2a02e1a6d2f92ee5469d083a3a1d8db03fc72d530d

Observation 5c6e3624-1cf6-460d-a82e-630435d229c8 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.807650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.807650Z digest=sha256:1f78e4271b41f0bd57b03178c52fa5ea20ff08c5d31426a4109f508df9dad5ee

Observation f52b7b31-46da-41ec-815f-3d1ac87b91e2 · outbound

This paper cites Training language models to follow instructions with human feedback.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Training language models to follow instructions with human feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.811437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.811437Z digest=sha256:df8525e383a8d3669dc0babc9f29c7c951dcea28bfce1f80777476ae8d85eee3

Observation e9d98946-ae4a-4eb0-acfd-02da1a80b42a · outbound

This paper cites LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.814907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.814907Z digest=sha256:85b2352c372180b07afc13c768c3c70741e156e4a3e901bb96f2927c7b1b6ff9

Observation ea4c7205-f1b4-4f54-b765-64f8e0121f7e · outbound

This paper cites Universal Jailbreak Backdoors from Poisoned Human Feedback.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.818464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.818464Z digest=sha256:22ae9c3884c88eb6a85f82b5bdd10832c81e8813aac44f6c0ec3cdf80f6963c2

Observation 8b8d9d38-736e-4310-821b-d742885177bc · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.822255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.822255Z digest=sha256:6de5499d0f9456e9f4e8dfbc5d0e0035a71c6285ba92c7bd8d826d88df608ff1

Observation 4146389d-132b-4bd6-8081-1fe955e307fc · outbound

This paper cites CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.825816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.825816Z digest=sha256:a64443dd1d552568b0972ff5773930bcaa88e096384e07ae922685542389ba62

Observation 44356f20-045e-4da2-9e84-8b14a768a0dc · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.829433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.829433Z digest=sha256:f6b62eb27ecc65e320afd3b0ddbab48d8f7d2ef0e6cd0813a8ede4aa3cbcb050

Observation 33988b2a-c35b-48c8-b3ec-b58bdf2dea45 · outbound

This paper cites do anything now.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning do anything now

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.833240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.833240Z digest=sha256:f5c29c34046d2190d118d76340ace900029cfa4ed8a91f5f3eab2058dd90676e

Observation 7c72ef52-306c-49f2-815e-41f77be8cc23 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Qwen2.5: A party of foundation models, September 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.837049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.837049Z digest=sha256:74d9c63259fd3875abb1cd37fd93a151c14be690509c49adb698302fb3f89ad9

Observation ff03216b-92ee-4574-83e4-69f22ea01aba · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.840899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.840899Z digest=sha256:8772c01bc5e79ada5b7aa6e0125bad6a316a5cc51cfeba4c820dc3f47f9dd67f

Observation eeec823e-f6ed-45c5-9d6c-b5e2ec9ce37e · outbound

This paper cites OpenChat: Advancing Open-source Language Models with Mixed-Quality Data.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning OpenChat: Advancing Open-source Language Models with Mixed-Quality Data

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.844524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.844524Z digest=sha256:a8e48b4903c2e03bf7f6d4bc864bc3c619c558164c290d5bed9a70a93c451382

Observation 8a8eaf54-aad6-4d17-9a68-39953e2e431d · outbound

This paper cites Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.848329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.848329Z digest=sha256:cdac9532adf1079538aaef0d6e2e89cf339b68500fbfee49a509f5a544ba8e26

Observation 524a50a4-a35c-481a-ac54-2ececa3686b3 · outbound

This paper cites Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36, 2024.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.852850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.852850Z digest=sha256:472637457dcda312bc62a86e867843772bbf981c81db1bc2067d9e5c6a605a54

Observation ed581a41-2e19-4c3a-93a0-50d01a279b19 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.856473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.856473Z digest=sha256:4e46730461260dbb302808bf32678a850888216bb99ca97779a36ae1b22b32ce

Observation 1738ef74-4719-47c3-8ad8-c8494ed086b2 · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminder.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Defending chatgpt against jailbreak attack via self-reminder

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.477073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:52.860142Z digest=sha256:77d4ccb7891ae52e7da68c9f12eccb2e468ef4704d61f0ecb3264ae7d75dd48a

Observation 493da253-3d00-4cab-9e27-74a02732220c · outbound

This paper cites Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.863562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.863562Z digest=sha256:c0257f3e4eb948ac1c1d69246307e14d049f06499caad4fb55a5b4f3d6a01a7a

Observation 4383d3df-5b9c-40cb-9135-33ab224320d0 · outbound

This paper cites Qwen2 Technical Report.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Qwen2 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.867549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.867549Z digest=sha256:c195c11366e66fa867d950c09af3099d8d4837204710e1cc8056153abda8fa73

Observation 35f49797-2e6a-4e80-ad93-790aa4ed03df · outbound

This paper cites Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.870948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.870948Z digest=sha256:3f2352e8c90c28540811a912c84d298658a19fc2d0415f74ca8e51f00ba795a9

Observation 99b54b90-02ac-41d0-919d-823d6d7cfc14 · outbound

This paper cites A survey on large language model (llm) security and privacy: The good, the bad, and the ugly.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning A survey on large language model (llm) security and privacy: The good, the bad, and the ugly

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.874542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.874542Z digest=sha256:78264f028c18fb427ea0e0230ebb4e55481ef1b3d55c4ede53c93d5e0a956596

Observation e088c2cf-5dc4-4985-8771-c216233db363 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Low-Resource Languages Jailbreak GPT-4

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.877810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.877810Z digest=sha256:2eeecfbe0080ee54b75b40944cf36df7b3b3989e32ec9675cd149887caa82d3e

Observation 095ec9a2-ea40-4352-a2c7-27ea35e465f3 · outbound

This paper cites GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.881332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.881332Z digest=sha256:9ef09633ea6ea334f2ab7a33bee0bc9e0e678414376db59bf4873c92e6f3af37

Observation de2ff848-1cfe-45d1-91dd-fbcea379715e · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.885191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.885191Z digest=sha256:c35aef5b76652e86f7c909fa3a2b0ea39200ddf1ef3897584dd8769a2512c090

Observation 34f0c01b-75ba-442b-891f-23955059e5e6 · outbound

This paper cites Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.889218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.889218Z digest=sha256:52309c990260bd6f48c6a237f323996afc8cb50cc00743efe5bd8d01c8c11c2b

Observation 2424d0e2-b828-4159-bed6-f8e2cce7f211 · outbound

This paper cites Diversity Helps Jailbreak Large Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Diversity Helps Jailbreak Large Language Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:35:52.997180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:52.893040Z digest=sha256:4e7dc02937b505f21ce404e2159ca45985e6a4c2bf7997aa43c320e3c6945a27

Observation 1f9f5f83-9ceb-4d85-b6b7-63b0cdaa32b1 · outbound

This paper cites Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.898323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.898323Z digest=sha256:75601256a1feb5dbc9d51727f68a1f0b05d701c646b0f5e67664b42dfb2b5355

Observation ed8e2339-b36c-4fba-a62a-9510ab9119d9 · outbound

This paper cites Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.902122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.902122Z digest=sha256:b8afbe6767595e582593cb3174aec7b5e65b275342859ef39a8a89c1da1aeac4

Observation b9476d38-a2a4-4bdb-b25d-3c64abba7121 · outbound

This paper cites Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.455117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:52.906616Z digest=sha256:427a3233298b2d96deb41e8282d063a9c37d55886e6e04efd4b47d8ae477547e

Observation dea701a4-d2cc-4c57-96ba-73755f0076f1 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.910697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.910697Z digest=sha256:33c4c57f845ad0e1903a1de447abb6ef5508eec99131b3e14b734d21adfdc600

Observation 47c3d999-626e-41cc-8893-c678846a02e7 · outbound

This paper cites As shown in Table 10, our method can still achieve remarkable performance on HarmBench [28].

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning As shown in Table 10, our method can still achieve remarkable performance on HarmBench [28]

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.439187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:52.915101Z digest=sha256:529f6d04a64242b38fb87a3e92393789076c20d0ec7f33966a899fda8325ff6e

Pith citing papers

No inbound Pith citation observations are available.