Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:35:52.915101Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2501.07959.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:35:52.915101Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 51987e9b-db86-4e42-85d4-7ed965eb8c5d · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Llama 3 model card
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e399931-dc77-41bf-a8d0-8557088251a1 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Detecting Language Model Attacks with Perplexity
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3af79303-da0c-41c9-affa-cbf70b6b7e2c · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 094b98a7-515f-4d18-b11a-c305be6ca13e · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Many-shot jailbreaking
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6bb033e1-ce8e-48c9-a83d-809d6234bd9a · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Foundational Challenges in Assuring Alignment and Safety of Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b95c695c-de00-4f27-8d69-e1594b0a73d9 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb0f603f-c2e7-4512-bf45-3b8ccb3bf130 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Language models are few-shot learners
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62be13bd-c7b2-43f6-9867-100bd8c422fa · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddfa3eb9-4a42-48a3-abe4-935f61abb4e2 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6413099-72d0-4e21-b4d8-71bf0ae35d1e · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2138c5fa-c552-429f-a962-5fed6f1913ae · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Combating misinformation in the age of llms: Opportunities and challenges
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d825ee4-f14e-40cc-a7dd-8430b73b6ff0 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Safe RLHF: Safe Reinforcement Learning from Human Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a5e26a7-1a84-457a-9088-1aecc83d8781 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Multilingual Jailbreak Challenges in Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a878be1-1938-426c-8f49-4cc7ed5d7ebb · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc414b65-0879-4503-8536-00c59249ef86 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning The Llama 3 Herd of Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc2925ad-506c-4846-a1f8-ed75d3afb945 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Red- teaming for generative ai: Silver bullet or security theater? In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pages 421–437, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 86a9f30c-9aee-4579-b0b6-5d765cd04cf8 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Badllama: cheaply removing safety fine-tuning from llama 2-chat 13b, 2024
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4aebbcbf-ea7f-4815-80f5-81b755a6aee6 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a230f98-a27f-41a7-9c84-bd66a97c4009 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 641c1b60-bddc-48c9-aeeb-6d8c3eeda8db · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Perplexity—a measure of the difficulty of speech recognition tasks
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b52040c0-11c3-4c80-9989-2315068e87fa · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Mistral 7B
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1af48f6-8b91-426a-b038-02fffdf6b462 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5580c03-ad9d-4d82-a648-b82d36d9a824 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d31ed7c8-b58b-43ea-af7d-f97174656187 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e59b2646-f26c-4178-a01d-b79d06bbc569 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning DeepInception: Hypnotize Large Language Model to Be Jailbreaker
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 774b9cd4-2bb3-44be-9d77-7e489c0def0f · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning RAIN: Your Language Models Can Align Themselves without Finetuning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d3a996f-7d66-422b-affe-4e8ac9491e96 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c6e3624-1cf6-460d-a82e-630435d229c8 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f52b7b31-46da-41ec-815f-3d1ac87b91e2 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Training language models to follow instructions with human feedback
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9d98946-ae4a-4eb0-acfd-02da1a80b42a · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea4c7205-f1b4-4f54-b765-64f8e0121f7e · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Universal Jailbreak Backdoors from Poisoned Human Feedback
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b8d9d38-736e-4310-821b-d742885177bc · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4146389d-132b-4bd6-8081-1fe955e307fc · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44356f20-045e-4da2-9e84-8b14a768a0dc · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33988b2a-c35b-48c8-b3ec-b58bdf2dea45 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning do anything now
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c72ef52-306c-49f2-815e-41f77be8cc23 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Qwen2.5: A party of foundation models, September 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff03216b-92ee-4574-83e4-69f22ea01aba · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eeec823e-f6ed-45c5-9d6c-b5e2ec9ce37e · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning OpenChat: Advancing Open-source Language Models with Mixed-Quality Data
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a8eaf54-aad6-4d17-9a68-39953e2e431d · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 524a50a4-a35c-481a-ac54-2ececa3686b3 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36, 2024
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed581a41-2e19-4c3a-93a0-50d01a279b19 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1738ef74-4719-47c3-8ad8-c8494ed086b2 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Defending chatgpt against jailbreak attack via self-reminder
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 493da253-3d00-4cab-9e27-74a02732220c · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4383d3df-5b9c-40cb-9135-33ab224320d0 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Qwen2 Technical Report
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35f49797-2e6a-4e80-ad93-790aa4ed03df · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99b54b90-02ac-41d0-919d-823d6d7cfc14 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning A survey on large language model (llm) security and privacy: The good, the bad, and the ugly
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e088c2cf-5dc4-4985-8771-c216233db363 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Low-Resource Languages Jailbreak GPT-4
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 095ec9a2-ea40-4352-a2c7-27ea35e465f3 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de2ff848-1cfe-45d1-91dd-fbcea379715e · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34f0c01b-75ba-442b-891f-23955059e5e6 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2424d0e2-b828-4159-bed6-f8e2cce7f211 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Diversity Helps Jailbreak Large Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1f9f5f83-9ceb-4d85-b6b7-63b0cdaa32b1 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed8e2339-b36c-4fba-a62a-9510ab9119d9 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9476d38-a2a4-4bdb-b25d-3c64abba7121 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dea701a4-d2cc-4c57-96ba-73755f0076f1 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47c3d999-626e-41cc-8893-c678846a02e7 · outbound
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning As shown in Table 10, our method can still achieve remarkable performance on HarmBench [28]
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
No inbound Pith citation observations are available.