Pith. sign in

Paper Citation Record · LEDGER

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models

As of 16 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 8 inbound Pith citation observations for arXiv:2501.01830.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.01830 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:25:54.739885Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:31:17.631468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T13:35:46.511843Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9e845f25-ecb1-428e-bf8b-e8cfb61fd540 · outbound

This paper cites write newline.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.517346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.517346Z digest=sha256:5e127e31ea309054caf336ff69b58f3d7251b92e3056a5bcaa6b8501acc2e02b

Observation 8c3c5024-00f3-40ad-b79e-84faf57cba39 · outbound

This paper cites Constrained policy optimization.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Constrained policy optimization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.523608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.523608Z digest=sha256:967f4171cdcb7e0cdd71ea680d33913c2bdc7eddb7c9d1d068bb14d35afa5b23

Observation 798bce14-0b0d-4892-adcb-37ed42899ca8 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Yi: Open Foundation Models by 01.AI

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.528176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.528176Z digest=sha256:b8c520edb75fbecaced6ddda98d849db22919c28cbc406b16fead275d462fbe9

Observation e2dbb23a-3724-460b-b169-8a805c30e9ea · outbound

This paper cites and Cook, R.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models and Cook, R

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:25:55.520386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T22:25:54.533228Z digest=sha256:ded3d7efb6faa60ccf878f0b4e48ad0bbb1f28d6569bb5f835c787f8bb17b282

Observation d679dd21-39bb-4686-a3f8-d785b220b1b9 · outbound

This paper cites Constrained markov decision processes.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Constrained markov decision processes

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:25:55.505394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T22:25:54.537586Z digest=sha256:df6a43a423c2853b8b6d09df0e3ba556b7f8178b07b8adf07ce257e858631817

Observation dbb8623c-7aa3-46b4-9cd5-b4a19c048a0a · outbound

This paper cites Does Refusal Training in LLMs Generalize to the Past Tense?.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Does Refusal Training in LLMs Generalize to the Past Tense?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.542105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.542105Z digest=sha256:0312d549aed40cf8ca5df3ffb0f189346917ea3018b9b9ee1cb79d45f4ca9648

Observation 6afc68dc-c61f-46b8-9a14-16ae69c059c8 · outbound

This paper cites an unresolved cited work.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.546943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.546943Z digest=sha256:a37dae12647191ec67b18c8b34f40a1176473d92381503a8d04451749a8a4b59

Observation a8f4a11f-591b-4bc8-8177-9b4103c39a26 · outbound

This paper cites an unresolved cited work.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:25:55.483143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T22:25:54.551471Z digest=sha256:eeb79e8fb620f1d67a68576c3a57846a4b502845027f19c6536385c004ce4f30

Observation 5c69d931-73a6-4970-a13d-9ff1d4eacbfe · outbound

This paper cites and Bailey, D.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models and Bailey, D

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:25:55.469140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T22:25:54.555359Z digest=sha256:11ab3828cdc0fb24af4b9cd4cf6a72d2599f2d8a0d8b9f1e63805694e443bfe3

Observation 521d5ff0-48c6-4870-ab52-1517ec408806 · outbound

This paper cites an unresolved cited work.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:25:55.455142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T22:25:54.559580Z digest=sha256:646cea1a9a977b57f629ae48c6c98327301fc07ef1b01d8ea848f855ab1f0236

Observation c0c23b09-df15-4136-87ed-639ce2894606 · outbound

This paper cites K., Savage, S., and Voelker, G.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models K., Savage, S., and Voelker, G

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:25:55.439927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T22:25:54.563620Z digest=sha256:fe20efebaa0e14d77a70512f57bb6dcd35e15f5dbe36e8a9cab4f157b77a59a3

Observation ce11943d-57af-4f62-a2ca-6367b14e4abf · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.567857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.567857Z digest=sha256:8910c006d6145b5aec5624dd8c0ffe25b8dda249a58fe82a361c08f2424246fe

Observation b32a6893-a88d-4b45-95ae-ecef7063250d · outbound

This paper cites E., Stoica, I., and Xing, E.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models E., Stoica, I., and Xing, E

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.572406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.572406Z digest=sha256:5e1303af8ea7a0fdd9b0ab2c56d0150ce079df709e3b8a834701ffe92eb08eed

Observation 3bcaff9c-3acd-4174-8fba-34545924e734 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.576434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.576434Z digest=sha256:bbc82e9e774f2cc61889b7b53fe42580be7557532bc736b8dbe6eeddf3075131

Observation 4c0012d2-f49d-4bae-a40b-fe3ad1d8c73c · outbound

This paper cites The Llama 3 Herd of Models.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.580690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.580690Z digest=sha256:096a25c4bef6cb833e86e7cc40ae50c3ee327678d5f4b8533865b5cb91090695

Observation c14162d4-1025-4355-86fb-efc46394a0f1 · outbound

This paper cites Challenges of Real-World Reinforcement Learning.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Challenges of Real-World Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.585351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.585351Z digest=sha256:2d496f6ec8c041a48edfbb1a03ee55c1101323dd56b2fcc66481a563548c0b73

Observation 7dff4194-51d3-481f-9f1e-907f38d1b702 · outbound

This paper cites MART: Improving LLM Safety with Multi-round Automatic Red-Teaming.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models MART: Improving LLM Safety with Multi-round Automatic Red-Teaming

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.589698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.589698Z digest=sha256:12e2820aaeebb70214f08352b93351f7c75ea49922a6744899f396a70850f0f8

Observation a8d92230-73ab-4d7f-842e-27320e4e435e · outbound

This paper cites Gradient-based Adversarial Attacks against Text Transformers.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Gradient-based Adversarial Attacks against Text Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.594335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.594335Z digest=sha256:5a6163722a4029609927954c449935e7f0b8d8a654ae5c66001d76570aa2339d

Observation 09129e28-8378-4284-81ed-04b0ac547eba · outbound

This paper cites COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.598462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.598462Z digest=sha256:5946513a1598b2558fe4a2635b3051605ce8cb0c636faa8cbb1fa7073d7db70e

Observation b9b4af55-9ea7-42cb-adb2-e89b78c9a7c7 · outbound

This paper cites Curiosity-driven red-teaming for large language models.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Curiosity-driven red-teaming for large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:25:55.417831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T22:25:54.602587Z digest=sha256:68e7174e61a8d232028e4f3a55a266c1e16d919053ef5a98aa4d14fa8fdb8332

Observation 2db33961-23dd-4b42-968b-3eb62241ab12 · outbound

This paper cites Mistral 7B.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.606381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.606381Z digest=sha256:03245e45eacb620ccc61734f686a2d7d137812d755273dd8b610b99283cec4d1

Observation a0ea4762-ffe1-460a-b9f9-87626d67ea2c · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.610490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.610490Z digest=sha256:af7bf404012e44e4baf9083d48bc2380aef66ae4b79559beed02214c8718a695

Observation 30900c49-d046-47af-951d-ce9610ed1ad3 · outbound

This paper cites AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.614748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.614748Z digest=sha256:b29f90938da866aab9f4a12c52543479dd67c73b6736d9ae9807e10bcca1bf63

Observation 92116db0-d5b0-4e42-9e39-e6f23640ae46 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.618480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.618480Z digest=sha256:e16b03581057a11be2bb3598c900253447c87bd788440e2628a6fd7104dda6a4

Observation bea2d9d4-7d88-4146-aa9d-cedd59fb6d4f · outbound

This paper cites Summary of chatgpt-related research and perspective towards the future of large language models.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Summary of chatgpt-related research and perspective towards the future of large language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.622364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.622364Z digest=sha256:243acb9534a1f95d4b51b4fa50ab72aefa9874184b290344af387553d3fa22b0

Observation 073d849e-b3b5-4fd2-aeab-9f9812d9b6c0 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.626108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.626108Z digest=sha256:962e5eca23fdc3b92842c055510c935245dbf94027ab050e802bfafae8bfe6cc

Observation 6bce1429-58d2-491f-9b37-af18e798a38c · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.629811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.629811Z digest=sha256:a902f31f3553f9d3622534f5f421acc784d5dde9e01552d51e350b2f8dfe9076

Observation ad92feb0-3244-4c29-ac55-e8b1f6d86561 · outbound

This paper cites Llama guard 2 | model cards and prompt formats, 2024.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Llama guard 2 | model cards and prompt formats, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:25:55.404103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T22:25:54.633549Z digest=sha256:4eb70059dc9df93a47e92b1b9d191e7ca821cf1dc3a7370c0284c918eb415254

Observation 4ee6e890-3707-4d2a-97bc-e159ea09e38c · outbound

This paper cites Confronting Reward Model Overoptimization with Constrained RLHF.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Confronting Reward Model Overoptimization with Constrained RLHF

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.637458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.637458Z digest=sha256:8fd484e0d7d5a0397f1a6e2df98419e18dd9a21a332e64a8f2f32d4fca16f792

Observation 6d6f5143-5a1a-477c-9f37-550e43f7c0c0 · outbound

This paper cites Y., Harada, D., and Russell, S.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Y., Harada, D., and Russell, S

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:25:55.390523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T22:25:54.641722Z digest=sha256:03d97dfdf3c29cdfbfdce3dc18dc889d85fd83bc08a02d070d62773fb90c456d

Observation 80a06231-a5e6-408d-9827-1f7dd82d3ddf · outbound

This paper cites Progressive Safeguards for Safe and Model-Agnostic Reinforcement Learning.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Progressive Safeguards for Safe and Model-Agnostic Reinforcement Learning

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:25:55.034791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T22:25:54.645755Z digest=sha256:b34a0c8afec6721f72bdddec369c3959f04be8826a6d68482bca297abbe632de

Observation e2c2213d-6dac-459b-93f4-fb6c9dee3adb · outbound

This paper cites Training language models to follow instructions with human feedback.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Training language models to follow instructions with human feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.650079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.650079Z digest=sha256:0811bc372c570b3a6722f359f9e859735ced3a4b24de219e8f4eb2940d09bbeb

Observation d9476fc3-7504-46bb-9d58-78762d1b8bc1 · outbound

This paper cites Red Teaming Language Models with Language Models.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Red Teaming Language Models with Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.654401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.654401Z digest=sha256:598afed65b47358a96e870cd4511916a0fe6d1e17f70ac94ce50fcac570be400

Observation f826c524-d33b-42b0-b6ae-e216afddb3ec · outbound

This paper cites Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.659473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.659473Z digest=sha256:4bb56245a3dee051f02c154e35726e80d07b6503ba7eacda2817826bb5b008e5

Observation eaa6c832-81be-4895-8a91-dfc5d842476d · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.663803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.663803Z digest=sha256:ab475e6f2653031b23e728d3fac53071663a78db459560c9fb8adb3c2325cb1f

Observation 550403b8-cf75-4ab9-8269-7248486d1546 · outbound

This paper cites Reinforcement Learning with Sparse Rewards using Guidance from Offline Demonstration.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Reinforcement Learning with Sparse Rewards using Guidance from Offline Demonstration

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.668146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.668146Z digest=sha256:63a396d24b8b617843f27a38ad4d2f315b2eb74ef480a5fef5ae8a6efe919957

Observation dd77d7f6-02ec-4205-8652-27790d29f87f · outbound

This paper cites Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.672598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.672598Z digest=sha256:e36fa4d5d98ebe8bd1902722b988cc373e52e5b142b3005cdf19d7103049523e

Observation 0315ce1b-7d2a-4f4a-bfce-19e19a41e7f7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Proximal Policy Optimization Algorithms

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.676911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.676911Z digest=sha256:7504507370f9af710e7340ba51618b5c469f9f9119a08bd25f1bda8577691591

Observation fe7272cc-c41b-4c30-969b-76a4876ea625 · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.681275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.681275Z digest=sha256:8b5f5176fd6a570928cfc4af42a56570248c367d61415da9eb02f1806b0e15a9

Observation bdf49c32-6e8c-41cb-b37a-63fb5ef6cb84 · outbound

This paper cites AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.685378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.685378Z digest=sha256:449306e0bfa84dda76820375a673c1357c50683ff0be50a94d4855f2a9a48264

Observation 32738862-ddf9-48b6-9a89-e9cb41d4b104 · outbound

This paper cites Safe Exploration by Solving Early Terminated MDP.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Safe Exploration by Solving Early Terminated MDP

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.689496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.689496Z digest=sha256:953ef75cd4e7e4bc6db9c04047c374f7fe03ba5bce3e2d5fc2beb92821f19063

Observation 479ea283-8c1d-410a-bb8a-13a388053e3d · outbound

This paper cites an unresolved cited work.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.693893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.693893Z digest=sha256:ca79909a8256b551ba606088b97236e9863b5c0d60979258589b04d11a2b6b7c

Observation b2569346-1b70-414f-878c-29b61a9617cc · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Gemma 2: Improving Open Language Models at a Practical Size

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.697793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.697793Z digest=sha256:73d2ac926651816538ac0e6765d5d6767b01384a0633e85ea549be3fe88f7249

Observation 4029d34d-89d5-450b-9cdc-7952b4876c86 · outbound

This paper cites Introducing qwen1.5, February 2024 a.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Introducing qwen1.5, February 2024 a

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:25:55.367954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T22:25:54.702261Z digest=sha256:864e77a87c2cba7bb6a9da8846ee9c0630268238fa1636a832b5c80442289b66

Observation 692e3707-cf9c-4a5d-9b3f-027d9131e0cf · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024 b.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Qwen2.5: A party of foundation models, September 2024 b

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.706303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.706303Z digest=sha256:908180c76ff76723c8b7c9900fa357c523f7b5b8f29b10242d38de33b004d039

Observation cefe18f5-dce9-4093-93f0-f8042d9544df · outbound

This paper cites Evaluating the Evaluation of Diversity in Natural Language Generation.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Evaluating the Evaluation of Diversity in Natural Language Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.710403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.710403Z digest=sha256:1ebeac4d0ba80b15d9024829133ff663e15e2dc6ea25b9e06c9770ae471d6626

Observation f270f410-fcbe-4947-92c6-6ff9dae2b282 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.714466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.714466Z digest=sha256:c5ae8fa06b53114af5544a461426f0ea7b704e5bf4fc864ca5d39420f5b8eaa7

Observation 7cb9c10c-5dc6-46f0-a20a-d7a3b3c5e334 · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Zephyr: Direct Distillation of LM Alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.718218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.718218Z digest=sha256:9d9bcc47a0db204a0894ec27fdc969438b40ccdfef6082bc456c8f1bcfed35da

Observation 962b564d-6ac9-41c1-8db9-b5cc36a03762 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.723257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.723257Z digest=sha256:d17aeddafc6ca82595ff7c91a1d21f12e060d3b350caf6921c9b6657436c54b7

Observation 635fe8bb-f7b9-4652-83da-4eee64fdb1a3 · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.727702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.727702Z digest=sha256:ff483d97a29f6576f60cd64012fd17a2abc1809ed5dc78c0f69e67ef8db6e4a4

Observation 49a057db-d875-41b0-9169-d4bd65c026ff · outbound

This paper cites DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.731944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.731944Z digest=sha256:391cf6484aa9129e73f6663af384553cbf0a1115b904ef962c461245390ae409

Observation 6df875c8-52c5-4bb8-91f3-6c76500f792e · outbound

This paper cites Purple-teaming LLMs with Adversarial Defender Training.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Purple-teaming LLMs with Adversarial Defender Training

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.735983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.735983Z digest=sha256:40bdcd68c8da1e4b0b415fc8338eda5f4b741752aa7922ff00ec91799f3dde93

Observation 16d9dbb0-fdd1-4a1a-af52-3026d4b2d7b3 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.739885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.739885Z digest=sha256:1fc0e8a853223e03c4b7067ef31e7e7ffa009cb221bf35e132b779df6dd73663

Pith citing papers

Observation a031d7ea-8d7d-4aba-b5f7-4741d1b3ae63 · inbound

DeepRAG: Thinking to Retrieve Step by Step for Large Language Models cites this paper.

DeepRAG: Thinking to Retrieve Step by Step for Large Language Models Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T16:31:17.631468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:31:17.631468Z digest=sha256:67bca2077106f16b01649eda00f8fe761ed6d1de302bf9651b9333834bc762e7

Observation d0d40a5b-e08c-4b30-abcc-3a8f1b8ef07c · inbound

Lifelong Safety Alignment for Language Models cites this paper.

Lifelong Safety Alignment for Language Models Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.780943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.780943Z digest=sha256:c1118ac422726b18543339d6a4fbfa747aae38e817ec4c77d781f63b27b1a0f8

Observation 7ef22d1b-236d-4d76-b9a5-70094e50e563 · inbound

Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning cites this paper.

Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:09.082636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:09.082636Z digest=sha256:9e12112303add9c62d16ceb0bd25d107b95694b191e358ac3509a2d739d1dedc

Observation c72098f9-5c78-41b4-9fcf-a7936fadf178 · inbound

Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models cites this paper.

Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:06.366957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:06.366957Z digest=sha256:e8d2113c7a8fcf4b9ea3c192294fe044007aa3d18c8ce13c3839c7e93234f45a

Observation e13f4df1-45cd-439e-ac90-5f8b31b6071e · inbound

SoK: Robustness in Large Language Models against Jailbreak Attacks cites this paper.

SoK: Robustness in Large Language Models against Jailbreak Attacks Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:01:08.754558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T16:42:41.137808Z digest=sha256:117f9c5b5c7cff6e7ac6f9d1bb384cc9cbab298a6aae10188ae40922613ac839

Observation 0406f2c3-d165-41d2-a9e0-e8858dc22804 · inbound

A Systematic Investigation of RL-Jailbreaking in LLMs cites this paper.

A Systematic Investigation of RL-Jailbreaking in LLMs Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:40:52.581648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-11T01:32:42.151644Z digest=sha256:badb5b9679ec753a551b415e2c2d41e9978383b5f0e2a257141d0e4b0453097f

Observation e46035ef-e850-464c-b33b-4ba0161d5f47 · inbound

A Systematic Investigation of RL-Jailbreaking in LLMs cites this paper.

A Systematic Investigation of RL-Jailbreaking in LLMs Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.513609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:59:07.861941Z digest=sha256:89895a3236162adf35845623203ea1a9bea3f8e92661a4dabf1d55995ff96903

Observation 0437a9e8-bf6f-4e51-ab31-df5006ffa24b · inbound

A Systematic Investigation of RL-Jailbreaking in LLMs cites this paper.

A Systematic Investigation of RL-Jailbreaking in LLMs Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T14:41:06.420030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:41:06.420030Z digest=sha256:ab0797179a79c62e40803e4944e2d915555bff3d2f7fd7abd573dd1160ad2d9e