Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:25:54.739885Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 8 inbound Pith citation observations for arXiv:2501.01830.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:25:54.739885Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T16:31:17.631468Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T13:35:46.511843Z
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9e845f25-ecb1-428e-bf8b-e8cfb61fd540 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c3c5024-00f3-40ad-b79e-84faf57cba39 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Constrained policy optimization
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 798bce14-0b0d-4892-adcb-37ed42899ca8 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Yi: Open Foundation Models by 01.AI
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2dbb23a-3724-460b-b169-8a805c30e9ea · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models and Cook, R
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d679dd21-39bb-4686-a3f8-d785b220b1b9 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Constrained markov decision processes
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dbb8623c-7aa3-46b4-9cd5-b4a19c048a0a · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Does Refusal Training in LLMs Generalize to the Past Tense?
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6afc68dc-c61f-46b8-9a14-16ae69c059c8 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8f4a11f-591b-4bc8-8177-9b4103c39a26 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5c69d931-73a6-4970-a13d-9ff1d4eacbfe · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models and Bailey, D
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 521d5ff0-48c6-4870-ab52-1517ec408806 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c0c23b09-df15-4136-87ed-639ce2894606 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models K., Savage, S., and Voelker, G
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ce11943d-57af-4f62-a2ca-6367b14e4abf · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b32a6893-a88d-4b45-95ae-ecef7063250d · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models E., Stoica, I., and Xing, E
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bcaff9c-3acd-4174-8fba-34545924e734 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Safe RLHF: Safe Reinforcement Learning from Human Feedback
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c0012d2-f49d-4bae-a40b-fe3ad1d8c73c · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models The Llama 3 Herd of Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c14162d4-1025-4355-86fb-efc46394a0f1 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Challenges of Real-World Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dff4194-51d3-481f-9f1e-907f38d1b702 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models MART: Improving LLM Safety with Multi-round Automatic Red-Teaming
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8d92230-73ab-4d7f-842e-27320e4e435e · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Gradient-based Adversarial Attacks against Text Transformers
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09129e28-8378-4284-81ed-04b0ac547eba · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9b4af55-9ea7-42cb-adb2-e89b78c9a7c7 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Curiosity-driven red-teaming for large language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2db33961-23dd-4b42-968b-3eb62241ab12 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Mistral 7B
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0ea4762-ffe1-460a-b9f9-87626d67ea2c · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30900c49-d046-47af-951d-ce9610ed1ad3 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92116db0-d5b0-4e42-9e39-e6f23640ae46 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bea2d9d4-7d88-4146-aa9d-cedd59fb6d4f · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Summary of chatgpt-related research and perspective towards the future of large language models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 073d849e-b3b5-4fd2-aeab-9f9812d9b6c0 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bce1429-58d2-491f-9b37-af18e798a38c · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad92feb0-3244-4c29-ac55-e8b1f6d86561 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Llama guard 2 | model cards and prompt formats, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4ee6e890-3707-4d2a-97bc-e159ea09e38c · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Confronting Reward Model Overoptimization with Constrained RLHF
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d6f5143-5a1a-477c-9f37-550e43f7c0c0 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Y., Harada, D., and Russell, S
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 80a06231-a5e6-408d-9827-1f7dd82d3ddf · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Progressive Safeguards for Safe and Model-Agnostic Reinforcement Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e2c2213d-6dac-459b-93f4-fb6c9dee3adb · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Training language models to follow instructions with human feedback
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9476fc3-7504-46bb-9d58-78762d1b8bc1 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Red Teaming Language Models with Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f826c524-d33b-42b0-b6ae-e216afddb3ec · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaa6c832-81be-4895-8a91-dfc5d842476d · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 550403b8-cf75-4ab9-8269-7248486d1546 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Reinforcement Learning with Sparse Rewards using Guidance from Offline Demonstration
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd77d7f6-02ec-4205-8652-27790d29f87f · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0315ce1b-7d2a-4f4a-bfce-19e19a41e7f7 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Proximal Policy Optimization Algorithms
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe7272cc-c41b-4c30-969b-76a4876ea625 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdf49c32-6e8c-41cb-b37a-63fb5ef6cb84 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32738862-ddf9-48b6-9a89-e9cb41d4b104 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Safe Exploration by Solving Early Terminated MDP
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 479ea283-8c1d-410a-bb8a-13a388053e3d · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2569346-1b70-414f-878c-29b61a9617cc · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Gemma 2: Improving Open Language Models at a Practical Size
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4029d34d-89d5-450b-9cdc-7952b4876c86 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Introducing qwen1.5, February 2024 a
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 692e3707-cf9c-4a5d-9b3f-027d9131e0cf · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Qwen2.5: A party of foundation models, September 2024 b
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cefe18f5-dce9-4093-93f0-f8042d9544df · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Evaluating the Evaluation of Diversity in Natural Language Generation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f270f410-fcbe-4947-92c6-6ff9dae2b282 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb9c10c-5dc6-46f0-a20a-d7a3b3c5e334 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Zephyr: Direct Distillation of LM Alignment
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 962b564d-6ac9-41c1-8db9-b5cc36a03762 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 635fe8bb-f7b9-4652-83da-4eee64fdb1a3 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49a057db-d875-41b0-9169-d4bd65c026ff · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6df875c8-52c5-4bb8-91f3-6c76500f792e · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Purple-teaming LLMs with Adversarial Defender Training
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16d9dbb0-fdd1-4a1a-af52-3026d4b2d7b3 · outbound
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a031d7ea-8d7d-4aba-b5f7-4741d1b3ae63 · inbound
DeepRAG: Thinking to Retrieve Step by Step for Large Language Models Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0d40a5b-e08c-4b30-abcc-3a8f1b8ef07c · inbound
Lifelong Safety Alignment for Language Models Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ef22d1b-236d-4d76-b9a5-70094e50e563 · inbound
Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c72098f9-5c78-41b4-9fcf-a7936fadf178 · inbound
Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e13f4df1-45cd-439e-ac90-5f8b31b6071e · inbound
SoK: Robustness in Large Language Models against Jailbreak Attacks Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0406f2c3-d165-41d2-a9e0-e8858dc22804 · inbound
A Systematic Investigation of RL-Jailbreaking in LLMs Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e46035ef-e850-464c-b33b-4ba0161d5f47 · inbound
A Systematic Investigation of RL-Jailbreaking in LLMs Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0437a9e8-bf6f-4e51-ab31-df5006ffa24b · inbound
A Systematic Investigation of RL-Jailbreaking in LLMs Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.