Pith. sign in

Paper Citation Record · LEDGER

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval

As of 19 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 3 inbound Pith citation observations for arXiv:2505.15753.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15753 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:48.828674Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:55:01.175611Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T11:37:15.922072Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99b7ab13-ce2b-460a-adf7-242f3bd5d271 · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Detecting Language Model Attacks with Perplexity

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.522821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.522821Z digest=sha256:5ea0fbb2c3f14d69a464df59ab1f13763b36edb50fea7031c9cc13c45f7e2500

Observation bbb68cd4-3423-4fb7-946d-878e2646fb5f · outbound

This paper cites RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.529201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.529201Z digest=sha256:2b02f79244f1f47a61093371ab1f4a273c8ca2cf41e4fc82644e60b5c49d2645

Observation c64ceb1b-9e7c-49c2-a707-4a12513aec4c · outbound

This paper cites Qwen technical report.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Qwen technical report

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:52.609923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:15:48.535290Z digest=sha256:3f110c1646de8425911aac4a896afbe9cd0fcf1109b9b9b3070beb1e41677e3a

Observation 886a8169-04ba-48e4-b2e7-e55a244fca48 · outbound

This paper cites Constitutional ai: Harmlessness from ai feedback, 2022.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Constitutional ai: Harmlessness from ai feedback, 2022

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.540149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.540149Z digest=sha256:58b79148c17d8f0e852b5125f8accea8c5e11a2fccd810f4741930ef6b2aa409

Observation b45cab46-1d09-4a98-879c-4da8a4448c56 · outbound

This paper cites Safeinfer: Context adaptive decoding time safety alignment for large language models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Safeinfer: Context adaptive decoding time safety alignment for large language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:52.384358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:15:48.544775Z digest=sha256:1d0d60af2a73d2c94bec554b1a66f52dbe39ac689e7d7a9bb8dc40e29ff4789c

Observation 98ec7c4c-dc42-4b15-8d62-997bdb9eb3a4 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.550190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.550190Z digest=sha256:183df96b69ad09b64691d6edf4dfd1ff5e2d89b0f21d312b0ac86d8a688a062c

Observation 5d460c8a-66b5-425c-ac4e-76febce85e8c · outbound

This paper cites Towards the worst-case robustness of large language models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Towards the worst-case robustness of large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.555459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.555459Z digest=sha256:09a31eb936e7eb6877fb9bfa273b62451738f6a9f04bb3177202e59ec00cb613

Observation d1712231-c99f-48e3-a1a8-e34fb7c9b811 · outbound

This paper cites Evaluating large language models trained on code, 2021.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Evaluating large language models trained on code, 2021

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:52.170794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:15:48.559854Z digest=sha256:245641e672b5e40beaee0a1a4381ce9d7f376c7958535c7755ffecd85bfa82d7

Observation 3db68c66-159c-4920-bc1a-6c8fe9a743c8 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Training Verifiers to Solve Math Word Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.565245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.565245Z digest=sha256:832f85c0564ff5f7edf6ce98949f925dd055427a2e2fb3db7eb507eacf40d2e6

Observation 00503b70-36cd-4723-bc38-02c821a144c7 · outbound

This paper cites Safe rlhf: Safe reinforcement learning from human feedback.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Safe rlhf: Safe reinforcement learning from human feedback

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:51.667541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:15:48.571116Z digest=sha256:bbc3790b7faeab1e26901ad85dee0e58c6f3968cee6edbb178c1f35ca3533a0b

Observation 316fc115-9d99-4232-b4e7-4d35eac70880 · outbound

This paper cites Multilingual Jailbreak Challenges in Large Language Models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Multilingual Jailbreak Challenges in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.577618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.577618Z digest=sha256:bb326a9424aa6e5669a45ab4dff460d41a4f3c9ef97756cb402c1989e3f9d4ac

Observation a248c4bd-a27b-4259-ab97-eee247e8dff8 · outbound

This paper cites A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.584004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.584004Z digest=sha256:8150193dcbf32f6e8c3e606f160e5b4e4c23a0eec300bd0adc6edff8a5cf63da

Observation 1e1a44d9-31bf-422e-b958-1cae8ab91900 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.589805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.589805Z digest=sha256:bd8957482597e4daf0b8fc15e4dc427c954341661a4744da191a289eaabc9e25

Observation f58024da-2a64-4c2a-adc5-c54e21b63d16 · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Explaining and Harnessing Adversarial Examples

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.594740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.594740Z digest=sha256:9da08c417331626954aca20ef0b91a101f853f841b363d81fced51f5f8357520

Observation 5c6df03e-c5fc-4ac7-808a-95433d683301 · outbound

This paper cites The Llama 3 Herd of Models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.601205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.601205Z digest=sha256:d81d0adc64b4244713ea482cd6b124a55e1fc8c0f2e3c68178413634f812efcc

Observation 5520535c-e1fb-4184-9649-3fd34870b845 · outbound

This paper cites Measuring massive multitask language understanding.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Measuring massive multitask language understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.605873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.605873Z digest=sha256:209cc8774e932ed4d58aed39ab5353bf9c1d55444355d79f84abba0ec745a34f

Observation 52d6f016-d02b-4072-953a-d445c93b68cd · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.609882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.609882Z digest=sha256:27462a968e844c446ffb5651066c2550e87f9ec1986b34f9228c9bd405be7284

Observation c8206a2c-c27e-4148-b506-cd9e25c144b1 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.615884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.615884Z digest=sha256:ce979640a621706d67bb819c2a0358fbb50bfef608ba7f7f970d334b1cb410e9

Observation f64e1a99-04dc-4715-9a5d-43ae87f3f4c1 · outbound

This paper cites AI Alignment: A Comprehensive Survey.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval AI Alignment: A Comprehensive Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.621390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.621390Z digest=sha256:56feb3bb198a16c1efb545a6b5c6b2c81400d5f0f7bf9cd83db76b0384abf339

Observation 8729fdfa-68d5-4471-96b8-893b157485b5 · outbound

This paper cites Improved Techniques for Optimization-Based Jailbreaking on Large Language Models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Improved Techniques for Optimization-Based Jailbreaking on Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.626254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.626254Z digest=sha256:f4f988f54ad7c2126b716f799118ac861af2017c047efdc2b92f77fee3b06b4f

Observation 1995933e-f583-4470-a0a7-35dbd5bd062b · outbound

This paper cites Mistral 7B.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.631589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.631589Z digest=sha256:281ac000eaf09d9980ba7e6662189af714217096e5bc483ddbc71346c97381b1

Observation 31bca02c-52af-4830-80f1-59a011fa3116 · outbound

This paper cites Artprompt: Ascii art-based jailbreak attacks against aligned llms.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Artprompt: Ascii art-based jailbreak attacks against aligned llms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.636714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.636714Z digest=sha256:3b193e663af05840b1c2176cb64649c3ab58f9ca40d25d20f3dbf515dee8539e

Observation 14b50c3f-e2d8-42a3-b948-a0a6529ea65f · outbound

This paper cites Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models, 2024.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.641406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.641406Z digest=sha256:0d7fe30f03d21bb2db803875ed677a3a4e7fb2b599179f79019baead2a8a749d

Observation b40304df-a0f8-432a-a08a-754a4a82db50 · outbound

This paper cites Dense passage retrieval for open-domain question answering.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Dense passage retrieval for open-domain question answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.646827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.646827Z digest=sha256:37553af4d629328b47860745b7d925b6f28ab5beb74c17354fc67ee82c4091cf

Observation 1c615d12-22ed-4350-9a2f-985701feaae0 · outbound

This paper cites Buckley, Jason Phang, Samuel R.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Buckley, Jason Phang, Samuel R

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:51.316907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:15:48.651724Z digest=sha256:accf93181ca4a54ab2ce4d8164e286045c646e080ed9ba47c970b064113d0a21

Observation 5df0e263-d0cd-4aa3-ae20-573e45d7e151 · outbound

This paper cites Open Sesame! Universal Black Box Jailbreaking of Large Language Models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.657664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.657664Z digest=sha256:29ad0f9d7e949fa0cce6b709b349708a078032aee01fe7fab36368adda3586e3

Observation f0e0be59-58ec-4d49-8f45-4ca9174e2b49 · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.663277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.663277Z digest=sha256:46fe0120ce668ca0e330b28933926b18ce9f1adb05f0209260e7d23a688f6e29

Observation 61a32c2e-917c-475f-8bbb-020ed343ffbf · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.669536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.669536Z digest=sha256:b1b694f3eb66a2de429f56bb0c800e6a37652bbc34031f80ab5c1728193134aa

Observation 97c73d0b-3953-49fd-9b60-d83085a05ea3 · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.674872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.674872Z digest=sha256:ec96a2fbd4dc3ba6a0ba42be3c5df82fa54c6b16538fca3b7fddf46ead97c570

Observation 21b586a7-03a1-453f-aca6-8674ddee56fb · outbound

This paper cites Autodan: Generating stealthy jailbreak prompts on aligned large language models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Autodan: Generating stealthy jailbreak prompts on aligned large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.679903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.679903Z digest=sha256:8b64f8d8edc2aa2995f96a88d4572b3bbb79d4dc1a181d795c4469e3e7583c7f

Observation 417498db-a5c1-4c68-af90-e1fe302d606b · outbound

This paper cites Jailbreaking chatgpt via prompt engineering: An empirical study, 2023.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Jailbreaking chatgpt via prompt engineering: An empirical study, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:51.112343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:15:48.685719Z digest=sha256:8b1bf4464fe61673591c8ef9eded4d4db0daf211cac5c447360bf9809f4a1535

Observation 3e2a2a90-9575-48b8-b1c4-1c77ed3fe5db · outbound

This paper cites Harmbench: A standardized evaluation framework for automated red teaming and robust refusal.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Harmbench: A standardized evaluation framework for automated red teaming and robust refusal

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.690350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.690350Z digest=sha256:da5463eecdd8b9acf1b70b9fb28c3a45c456433ef3afbad6b27ce792d7758a6b

Observation a42ad48c-47f4-4623-a68e-88ac9302a7cd · outbound

This paper cites Tree of attacks: Jailbreaking black-box llms automatically.Advances in Neural Information Processing Systems , 37:61065–61105, 2024.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Tree of attacks: Jailbreaking black-box llms automatically.Advances in Neural Information Processing Systems , 37:61065–61105, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:51.015657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:15:48.695248Z digest=sha256:32f270bd6f6f1acfcf08f4014e513a8fa7248a74e521915a9e3f4cbcd4571b78

Observation f9dd6051-07a6-4c12-9962-15a75bfd8fe9 · outbound

This paper cites Rapid Response: Mitigating LLM Jailbreaks with a Few Examples.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Rapid Response: Mitigating LLM Jailbreaks with a Few Examples

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.700337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.700337Z digest=sha256:c25f43a41b28e19d94345acd86ee4a4cfbf027192e834f9f7faece8a7b02a798

Observation 38744573-a46d-4814-aaf5-834a00156e4a · outbound

This paper cites Position: Adversarial ML for LLMs Is Not Making Any Progress.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Position: Adversarial ML for LLMs Is Not Making Any Progress

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.706076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.706076Z digest=sha256:be2c107d0b498d33685e4a26496ae22256cf3fc449e609026c8cba0f5a124421

Observation 13e1fa99-f5d6-49d1-923f-23fd1959b7b9 · outbound

This paper cites The probabilistic relevance framework: Bm25 and beyond.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval The probabilistic relevance framework: Bm25 and beyond

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:50.918748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:15:48.711127Z digest=sha256:fe0b202376a5d703065990c8806e6bc0709559ed2566e93455b45fdc627b5448

Observation 2a0ec857-ce8b-41a3-875c-f07ea6f12cb2 · outbound

This paper cites Mitigating skeleton key, a new type of generative ai jailbreak technique.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Mitigating skeleton key, a new type of generative ai jailbreak technique

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:50.810370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:15:48.716170Z digest=sha256:b615e198bb4ab94a755f8de428bb61be3ecc64c522f6b4ebf35b3837841694be

Observation f15bcb96-d1a4-4fa6-8051-2dbd064f51c2 · outbound

This paper cites Fine-tuning mistral 7b large language model for python query response and code generation: A parameter efficient approach.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Fine-tuning mistral 7b large language model for python query response and code generation: A parameter efficient approach

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.720684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.720684Z digest=sha256:0684410329d6a4407d7b76817c29d97aa86a0053893310c8c10937a6663de6f0

Observation b5bc4c4e-73aa-49d1-bfac-37998a2d099f · outbound

This paper cites Intriguing properties of neural networks.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Intriguing properties of neural networks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.725718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.725718Z digest=sha256:1802ba79a5fe153bbabdbb0a15813e025a63c63bc18abf8cf68995247e9f040f

Observation 07420015-32fa-4fbc-a748-10cfb72eab27 · outbound

This paper cites A theoretical understanding of self-correction through in-context alignment.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval A theoretical understanding of self-correction through in-context alignment

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:50.669636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:15:48.730652Z digest=sha256:1adcd83515179e5e9b48d6439351e6a28f54bd7f0b12a55b9498cebf62f32b65

Observation ed33eb3c-588d-40f4-b5d4-4f68de34dea2 · outbound

This paper cites Reinforcement Learning for LLM Post-Training: A Survey.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Reinforcement Learning for LLM Post-Training: A Survey

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.734701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.734701Z digest=sha256:c9bc008eb75d799d277153e561e48bfe8c162e294ca280714525c136f9149652

Observation bedce304-5b00-4bfa-a80b-78b18299dbf8 · outbound

This paper cites Jailbroken: How does llm safety training fail? In NeurIPS, 2023.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Jailbroken: How does llm safety training fail? In NeurIPS, 2023

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:50.569860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:15:48.740043Z digest=sha256:59474fe1d8d6e6ed1590e76607166c2f3c7877de42818924b64174003530bfbf

Observation 9bdd8198-b371-4ea6-b84c-bdc93100f652 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.745729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.745729Z digest=sha256:a5412fb4cfa6c2bce43ec3347c6d27a750a0d58258092beb3356fd98429e7657

Observation 80923eb6-3302-4e01-aefa-cd38a4aa9859 · outbound

This paper cites Certifiably robust rag against retrieval corruption.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Certifiably robust rag against retrieval corruption

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.750950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.750950Z digest=sha256:30eceb6c9d36c3f6f88e0f3ac1995f6c57c2a1de51a55e34b7086c7ed5aa343b

Observation f959c3cd-4489-4b6b-a410-054ac506d400 · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminders.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Defending chatgpt against jailbreak attack via self-reminders

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.756297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.756297Z digest=sha256:aec996df5d33ef42999227205080d8b1f2ee2ec9637b50df52a41ca5ce3af6b2

Observation 3f2115bb-ba62-431e-b74f-8fc169eb8c44 · outbound

This paper cites Safedecoding: Defending against jailbreak attacks via safety-aware decoding.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Safedecoding: Defending against jailbreak attacks via safety-aware decoding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:50.435836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:15:48.760766Z digest=sha256:a4b2bafec53e8dce075d5c9bbeac687baed9b57c9d86cbabad566b1c5ea14305

Observation d7afac52-f6a5-4c8d-ab70-f10d27949b6a · outbound

This paper cites BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.764754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.764754Z digest=sha256:96c2e987b03dfa8fd9a78513aaeb642e9c01d87f19f4d8c7c8e919126d2e59bc

Observation d2e0e766-d937-4532-ba55-fbcc41433d8d · outbound

This paper cites GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:50.261654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:15:48.770152Z digest=sha256:7bff74e53248116ab905c868be3e70f8d3923736ff6cab38a0c0014c18c91499

Observation 98157b43-ce31-4c47-8cb9-741cf18ad1b1 · outbound

This paper cites The ai alignment problem: why it is hard, and where to start.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval The ai alignment problem: why it is hard, and where to start

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.774898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.774898Z digest=sha256:85c0d013d0010c94de90d5ccb1ee71f2d00e82f82c6736b08651231dc76ff0ca

Observation 4648f707-0efa-4133-9d6f-04b7ed3ff604 · outbound

This paper cites How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:50.030436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:15:48.779396Z digest=sha256:99699266935d0b0f741ee283bc5efdef18a988df02fec677a0874f30f4c65145

Observation 80e84f86-572f-43b4-adf7-e09f88940881 · outbound

This paper cites Boosting jailbreak attack with momentum.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Boosting jailbreak attack with momentum

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:49.714847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:15:48.785059Z digest=sha256:e17c7560a9f4a278c0966f7e83ec7484c80d3ca257e23717f5baaa12dfd1b3eb

Observation 252c035a-fe45-41e2-80c2-4547273f67ef · outbound

This paper cites Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make your LLMs use External Data More Wisely.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make your LLMs use External Data More Wisely

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.790118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.790118Z digest=sha256:c9723619092db5a0e1332f966a174f72531383755f45b27220dec3edb8e5131f

Observation fde56ada-4e21-4d97-9e75-fc1a3582ac16 · outbound

This paper cites On prompt-driven safeguarding for large language models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval On prompt-driven safeguarding for large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:49.528671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:15:48.796220Z digest=sha256:cee1b0441c6f6e6a65d92421c074ce573788c7b28457d5440ee5983fccb296b1

Observation dee30cea-29b1-40b1-b489-8b9ad2672dad · outbound

This paper cites Poisoning Retrieval Corpora by Injecting Adversarial Passages.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Poisoning Retrieval Corpora by Injecting Adversarial Passages

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.802863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.802863Z digest=sha256:5f4b30552bebb21b2ebcc4088385deacc761e49775898ac83d77f6e1481224f9

Observation e4007808-2e3a-4b2b-9c68-e66246ff8219 · outbound

This paper cites Trustworthiness in Retrieval-Augmented Generation Systems: A Survey.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Trustworthiness in Retrieval-Augmented Generation Systems: A Survey

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.809235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.809235Z digest=sha256:0ba5a006d0b0cedf0e24a93e156b1ee6f93baa111a48ec26afc29055c5caa820

Observation 94337f36-ae35-4de6-9ad1-7e7532814868 · outbound

This paper cites ATM: Adversarial Tuning Multi-agent System Makes a Robust Retrieval-Augmented Generator.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval ATM: Adversarial Tuning Multi-agent System Makes a Robust Retrieval-Augmented Generator

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.816711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.816711Z digest=sha256:37d1dd23754c6e9570e975e0570537640e70b89400c4bc2e99dc51cbe35cfa46

Observation e7d43d4a-9670-4242-8b07-6e026f2d3d14 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.822543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.822543Z digest=sha256:c7326756e6dcb83943b00b1d85e852c9b0751c610ee47ee639dea51e06b5534e

Observation 26e778b4-da45-4896-8dfd-0f438f6e2eee · outbound

This paper cites PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.828674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.828674Z digest=sha256:0a6f10966a4bfa230d8885683af5b89fd15c4c7866cbb9f04a3de4a28ae8be19

Pith citing papers

Observation d0e4b5aa-4b40-4323-bfb8-6848d43a9507 · inbound

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction cites this paper.

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:37:15.923828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T11:34:09.428653Z digest=sha256:a29b994c715729e62b1adcdb3969210209c6c0b12db80c6cf1a4fd6608e28f93

Observation 8d7f6248-08d0-45b0-8e87-2bc492581934 · inbound

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring cites this paper.

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T22:31:19.445142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T22:28:41.134253Z digest=sha256:c15a6c50645b10909c5df218b5868577b6ee41c44cc15c85966b93cb9c68aed5

Observation a6d4da1b-bd9f-48fb-aa91-94f1b4938306 · inbound

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions cites this paper.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:01.175611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:01.175611Z digest=sha256:6487f7df9b74193da7bdd247f51a9b139da6e05f64951449b320736fa77de46f