Pith. sign in

Paper Citation Record · LEDGER

Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 80 inbound Pith citation observations for arXiv:2310.06387.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.06387 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 80 of 80 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:53:15.305158Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

21
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 117d23d0-55a9-4849-8078-a4f74db8fd9c · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.426486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:209c3ff1ec57cbf1b976ca9b03994b12be7364ad0cfc239c6ea8bf3985979011

Observation 9ca19ab0-3dea-4262-820a-f10d6891b276 · inbound

FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks cites this paper.

FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T18:41:06.153149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:41:06.153149Z digest=sha256:7f3bf6831d28cff259bffca1595dee0de686d92f61963f53759724597989d015

Observation c95fb98b-469c-4d12-a165-66b54dc5fc35 · inbound

Look Before You Leap: Enhancing Attention and Vigilance Regarding Harmful Content with GuidelineLLM cites this paper.

Look Before You Leap: Enhancing Attention and Vigilance Regarding Harmful Content with GuidelineLLM Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T18:53:15.305158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:53:15.305158Z digest=sha256:55c3b96d2aa06ea45e5b8c8a24d4c0e2a21de1aa0e5d83c136e41164773ea739

Observation 6fedc0ca-9974-49fb-8a0f-8024c1a93efe · inbound

No Free Lunch for Defending Against Prefilling Attack by In-Context Learning cites this paper.

No Free Lunch for Defending Against Prefilling Attack by In-Context Learning Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:49:38.142015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:49:38.142015Z digest=sha256:7816de1ed306bb28447b78dfb747797c214218902e92bdee75a54fe921529daa

Observation d24f0f6a-2382-49ba-b029-bfaf1f3e4515 · inbound

Towards Responsible Governing AI Proliferation cites this paper.

Towards Responsible Governing AI Proliferation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:48.601255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:48.601255Z digest=sha256:4c9a0c8782c7008689d8784e7b12402de31b3e7f2ea7dcc26d06de21139c6200

Observation 4cfd5c5e-551a-4de1-8066-5a1b12d2d127 · inbound

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models cites this paper.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.713244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.713244Z digest=sha256:e91e2cf631c9e765e9026a4223658883a049db93a4540f9883ed4353305ba8e6

Observation 9cc1274f-7c2e-475b-99c0-af3b7e4c34ac · inbound

LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models cites this paper.

LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:40:38.873180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:40:38.873180Z digest=sha256:d92c2f3a8d8216c796cbcaac8f4d88109cfdcdf0ceac6844cb23b8ca0cea6f14

Observation 7ef6edf3-7144-427a-8a56-67f480a69451 · inbound

SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation cites this paper.

SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:28:56.732300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:28:56.732300Z digest=sha256:5db369bc077330b0c0f81b8961c97c9fefdab17d5cca1b8389ab85cbd39d2e10

Observation 962b564d-6ac9-41c1-8db9-b5cc36a03762 · inbound

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models cites this paper.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.723257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.723257Z digest=sha256:163bac9dd8b36002240a40c790668069dd45c8a8a8c57a8b59501cf88d0d6641

Observation d6ef1e05-0fd2-4d45-bd4d-a8d6458c5f61 · inbound

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs cites this paper.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.984854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.984854Z digest=sha256:b3e4867ab5d08bf1f3fbec9c46cea6f47f6470fa25dfdf5e4c0dabffbfd332e1

Observation 6d605eed-3151-4592-a9e4-07be3a451d65 · inbound

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense cites this paper.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.912758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.912758Z digest=sha256:657dd05624ee6930d3fda5d99420edc2b06f9a0ed292ea7eb60c939609e15f8f

Observation ed581a41-2e19-4c3a-93a0-50d01a279b19 · inbound

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning cites this paper.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.856473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.856473Z digest=sha256:3eec301bb7325ed81188f1e7f65ceb656a00baf6fcb52174b0f5cbcd21870113

Observation 19cf3e0d-9de7-48fe-b988-f87b47517819 · inbound

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks cites this paper.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.737120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.737120Z digest=sha256:496d9e22b49a414073373c483f92bd9aa090c7f0ddfa2958ad1cbb05d506e8f1

Observation cb412c3f-1d20-49c5-8853-766be59699f9 · inbound

Episodic memory in AI agents poses risks that should be studied and mitigated cites this paper.

Episodic memory in AI agents poses risks that should be studied and mitigated Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T17:58:18.112923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:58:18.112923Z digest=sha256:b510028b9b4035641cd6f4fcd34ac6035457b32990458210b82094ec06a2f2b3

Observation 0175743d-88fe-4eb6-9851-69008c436cdf · inbound

PromptShield: Deployable Detection for Prompt Injection Attacks cites this paper.

PromptShield: Deployable Detection for Prompt Injection Attacks Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:10.811952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:10.811952Z digest=sha256:e0d553f2d9a7141b1dfdca42f1e4d3e5266f2739ecb0cb61416c2c70548d8a97

Observation 0f26dd78-4c35-4556-8f79-98f7892de47d · inbound

Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models cites this paper.

Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T00:12:04.817181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:12:04.817181Z digest=sha256:1f14d621415464841020378d3f2825783a1984425450f89a8549e7b03dbff04b

Observation a928bc4c-1dd2-43bd-862a-186b3966b4c8 · inbound

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling cites this paper.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.908552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.908552Z digest=sha256:f557a4f599a4afbbd6e674d63b00054a922b240fb79bd1bc8894e879f76fbc12

Observation e184b199-3148-4634-bc49-02b558a134dc · inbound

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation cites this paper.

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:30.673565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:30.673565Z digest=sha256:2e4ab1b1f1a36f8be95056a79f76f0959079e8e2627b758326985a24ff989e44

Observation e2a488f1-08c4-43a6-a2a4-94b88f78a080 · inbound

MetaSC: Test-Time Safety Specification Optimization for Language Models cites this paper.

MetaSC: Test-Time Safety Specification Optimization for Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:45.010325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:15:45.010325Z digest=sha256:c369eb801a243c42f1783a9061f629f3196ceda4865c195143ed1e428b7f6064

Observation cba2e2e6-8940-425b-aecc-e698ad06b693 · inbound

Advancing LLM Safe Alignment with Safety Representation Ranking cites this paper.

Advancing LLM Safe Alignment with Safety Representation Ranking Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.326956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.326956Z digest=sha256:f126a975f7378e432f933ffa53309a0a6f1b1efa8d21dff46a8ed12b9c5e37c6

Observation 9bdd8198-b371-4ea6-b84c-bdc93100f652 · inbound

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval cites this paper.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.745729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.745729Z digest=sha256:db8d9f4099e843892e5e818e22b867ffadbc9ed99f7ff006118a866c0ce9b059

Observation 380184e7-86ed-4d1d-acad-bd528028193c · inbound

Implicit Jailbreak Attacks via Cross-Modal Information Concealment on Vision-Language Models cites this paper.

Implicit Jailbreak Attacks via Cross-Modal Information Concealment on Vision-Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:41.028053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:03:41.028053Z digest=sha256:cbd5580438fc424a1d84d6898adde4dfb1687b2434fab698576daaaab5f6c7f0

Observation 6f79f5c5-dadf-42b5-871f-460fbc8e94b7 · inbound

Secure LLM Fine-Tuning via Safety-Aware Probing cites this paper.

Secure LLM Fine-Tuning via Safety-Aware Probing Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:11:35.853681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T13:07:09.402763Z digest=sha256:bc2b662d139c5e0f1422016a9f6cb1922ace5c7c79ba0d9b5ea276ed13a90c3b

Observation 50dc41cf-1536-4e7b-869a-64ad6d1d0615 · inbound

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation cites this paper.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.720380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.720380Z digest=sha256:9ceda1cf4f483e4fdf2ddd93d97785d850409f2d494703514212ec12187c6be6

Observation 315c024d-59f5-4a55-8fe7-56370521a7b3 · inbound

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts cites this paper.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:11.052058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:11.052058Z digest=sha256:f35ae04d25488e1a0a179df587ff2abaf5199c4d66394bfd55aa9281009e8a58

Observation 73d2de4a-8bd0-44b0-b6e4-7b1dec2610c8 · inbound

Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models cites this paper.

Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:24.102566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:17:24.102566Z digest=sha256:9f3e6b9a804bbb524724a0e0ade41b7ef6418c903138e93652d47d724daa8329

Observation 5faeabe1-ec16-46fd-8fbe-ec8a76ffd52e · inbound

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap cites this paper.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:22.236454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:22.236454Z digest=sha256:0924997c07457b3904c65c21f72d225492632ada32299c30e654d1aa0b2c63c3

Observation 75792518-7d97-4fe7-917b-fcd454ccace3 · inbound

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction cites this paper.

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:37:15.987956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T11:34:09.428653Z digest=sha256:1fe13301609f4b8f0fae8d937afba355ba87ee4672221cfd998dd216cb7a95df

Observation 60f67dd0-5532-4871-9215-078dcc65b23a · inbound

Adversarial Attacks on Robotic Vision Language Action Models cites this paper.

Adversarial Attacks on Robotic Vision Language Action Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:40.551096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:40.551096Z digest=sha256:bd98b9b81649ac6804c8ba414a0da1cc12db7a4819679b8da9fb1b32e9690f2a

Observation abd6921c-f91f-4c0e-b1af-47bf8baa381a · inbound

TwinBreak: Jailbreaking LLM Security Alignments based on Twin Prompts cites this paper.

TwinBreak: Jailbreaking LLM Security Alignments based on Twin Prompts Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:58.135736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:36:58.135736Z digest=sha256:35c4f84b9fb8408d9e10b6982cf447016df65703d0488e2602d216028b871b0e

Observation de60991d-b09f-4bc5-815e-e386c2095b22 · inbound

Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations cites this paper.

Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:57:16.542808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T10:54:55.262180Z digest=sha256:f90613ec38ba03f0b2a93b532f173c3d248318ffa44649588f9e97d6a5a8f7e0

Observation b255213c-9252-42fc-8015-5501bcc45f5a · inbound

LLMs Caught in the Crossfire: Malware Requests and Jailbreak Challenges cites this paper.

LLMs Caught in the Crossfire: Malware Requests and Jailbreak Challenges Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:10.112904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:35:10.112904Z digest=sha256:5c3eb678ae27abc7ea99136adc98ccc3ab08dd020b19ba0b9c3ac4c6fe6e427b

Observation 0bfa0aee-8356-4b43-a9f4-82adba2f8a86 · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:27.001406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:27.001406Z digest=sha256:529e63f1f5faae6172cb95f5b5179990d5b553f9d4127ba24f897ab7e7db0060

Observation 55f36b96-93fa-4d92-8919-e1cc47c74088 · inbound

InfoFlood: Jailbreaking Large Language Models with Information Overload cites this paper.

InfoFlood: Jailbreaking Large Language Models with Information Overload Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T01:02:31.037588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:02:31.037588Z digest=sha256:afaad0516ccd6851d2358af5ce43ff2fac545b1c422a71795bb429bf479c0940

Observation acd11a48-ed3f-4c4d-be94-32473105ddd4 · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:53.797609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:53.797609Z digest=sha256:17431d4c282e78587200947b90d2294f1ebbaccb6d1138a2e2a6784f193e2790

Observation c8564b34-057e-407c-bf81-6ed1135155b4 · inbound

Linearly Decoding Refused Knowledge in Aligned Language Models cites this paper.

Linearly Decoding Refused Knowledge in Aligned Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.014424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.014424Z digest=sha256:c62fecb01f32a4fedbdb7000c2661ade36247f02716b42c816114fb3663d4c39

Observation c9e12c66-942d-4f13-b196-686c4bf8756d · inbound

Defending Against Prompt Injection With a Few DefensiveTokens cites this paper.

Defending Against Prompt Injection With a Few DefensiveTokens Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:42.838856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:42.838856Z digest=sha256:5a3336e645f65cb219a93010bacffea546f85fd3d8b7cf05836a6997a519a3cd

Observation 78c43e03-bdbb-458f-83aa-5413134e33b9 · inbound

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation cites this paper.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.743717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.743717Z digest=sha256:a9c43a31bd3c07d225f41fd744acd5b8cf49175534589922f86a4bce83be8ea0

Observation bf2e2b86-141e-493b-a58e-0e75f00be4e4 · inbound

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models cites this paper.

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T16:25:55.100347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:25:55.100347Z digest=sha256:e3c337d8b9f5a7eaffaa4ff3a692cc7c5928d3b36806accaa642c3511b6dc055

Observation f0d19bb1-5cea-4a7d-928b-271ecf66ead2 · inbound

MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts? cites this paper.

MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts? Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:17:34.156403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:17:34.156403Z digest=sha256:97c2618fe831fb507eade88f1b8f0e96bef399cf907e2d084975d16b1ce53440

Observation 218856ef-8017-4313-b7d8-3c4288aa9d6f · inbound

PUZZLED: Jailbreaking LLMs through Word-Based Puzzles cites this paper.

PUZZLED: Jailbreaking LLMs through Word-Based Puzzles Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T05:43:39.293986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:43:39.293986Z digest=sha256:772d14b34ac1a8d10c8ab9540a413187bca7b5b9bee31f8d5d6b1368f2d3dd01

Observation b53d180c-a509-4bc2-8d06-8efb4e8c110b · inbound

Automatic LLM Red Teaming cites this paper.

Automatic LLM Red Teaming Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 1685

Resolution
unresolved
no resolver link, observed 2026-08-06T00:04:25.504135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:04:25.504135Z digest=sha256:fc44872ebbbf8616938f93dae462b168ae1cb00838431eb90fb06e44e90d91ea

Observation 11d2042c-6acd-4c2f-8ba3-21a2845c9a02 · inbound

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection cites this paper.

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T22:24:23.428703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:24:23.428703Z digest=sha256:aa8b5951fc80fa012c4b496239c602113106ce7a83f6fcac00846db0858811b2

Observation 7bbdb906-2757-4b09-abf5-0001b09d1ebc · inbound

A Survey on Training-free Alignment of Large Language Models cites this paper.

A Survey on Training-free Alignment of Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-05T21:18:48.551115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:18:48.551115Z digest=sha256:210a52f88a223aa52bd18cda58723af3153daf8822f98f2013b2c5d201a71491

Observation 3e071d78-c297-4a80-8084-43d8cac74c4f · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:41.559678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:41.559678Z digest=sha256:f5a86c9bf1fc4ba8a1ce3320c2bc64e662bc108100f01d44ca6901dde4a75647

Observation 103a5d3e-866c-42b7-a0b0-5c77d922d988 · inbound

SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks cites this paper.

SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T18:05:32.424178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:05:32.424178Z digest=sha256:31d7a17659a51bd0060203b1b00b69606a9e3e103d0768f819e511c533480090

Observation fe642ff8-1d94-4870-a75b-bc2df70e9cc0 · inbound

On Surjectivity of Neural Networks: Can you elicit any behavior from your model? cites this paper.

On Surjectivity of Neural Networks: Can you elicit any behavior from your model? Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-05T16:00:52.110346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:00:52.110346Z digest=sha256:4deb3d3073fe5082992ec46b4f7dc799a7349e18d06235e54453c5c7821a8cca

Observation 921a1417-8384-499b-97e9-4bdae34b8f08 · inbound

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs cites this paper.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:36:52.479691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:596462a92d549ff6fa2cb4c941a4290d35918ec8522cd02bfe665160e4aaefea

Observation 262a77df-7593-46d2-b169-d8fe64bae7d3 · inbound

Baichuan-M2: Scaling Medical Capability with Large Verifier System cites this paper.

Baichuan-M2: Scaling Medical Capability with Large Verifier System Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T11:50:14.742567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:50:14.742567Z digest=sha256:e8d6844939b79a4832b97b849305aa65a3b53f6353ce925574a7b18368b52dd0

Observation fcf06fb2-f84c-486c-9e38-4c9f4c0ca189 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 264

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:08.207569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:08.207569Z digest=sha256:42e71033ca79259e8d86b161cf09cb932b9051fbc40fbc83d15f474c84fd824b

Observation 881ed3e0-8875-48cf-a485-a295d796bf27 · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.470662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.470662Z digest=sha256:4d22f8f495024ca95981c1777d0b605d8332af8c3b01f516bedaf797fd2b389f

Observation bb991391-33fa-4d40-b1c4-1646f578c6a8 · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 200

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:57.778089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:57.778089Z digest=sha256:bdcaddc0eeec6beab7c142de73e9c30245ca362933efd40930333eeab24b7f22

Observation 8b9af65f-eef9-4e39-b148-2c7c4b2f4287 · inbound

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models cites this paper.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.630924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:1fa0396a454da99c25ea8db46797de565383b46bd303545daf176aa5fde6a09e

Observation 5187e3f8-ddbc-49ef-b93f-23f7013e2313 · inbound

Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels cites this paper.

Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T07:06:38.120540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:06:38.120540Z digest=sha256:51dc84bb1bb3c20b05973a9114b80d2c419a5396b8a52fc10299c73b85baec49

Observation e241d81d-697b-4658-86dc-2532dc2b8cd0 · inbound

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs cites this paper.

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:55:38.050506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T01:54:22.995178Z digest=sha256:37b00b121937c74156c7b1ba7683a066502c5b5a01a90b8e171852118ebbd753

Observation 1c861905-1bde-4f80-b222-a54f7756b3e8 · inbound

GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents cites this paper.

GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:05:26.677753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T07:02:28.660002Z digest=sha256:405cb022c4442b4c35e5bbb45e3b33f708677a2b6ac95894c01661fc96d1c1f6

Observation b6802b60-e140-435f-8723-813ea5ddc511 · inbound

ContextCov: Deriving and Enforcing Executable Constraints from Agent Instruction Files cites this paper.

ContextCov: Deriving and Enforcing Executable Constraints from Agent Instruction Files Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:50:12.375753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T17:49:08.192347Z digest=sha256:d7dcf0e51aea4a2c8a576e50da2924de279ee7752ce7fed649591dc71dee3a73

Observation b6f4ff22-d27a-4aa8-b762-1b9e0680ce30 · inbound

TrajGuard: Streaming Hidden-state Trajectory Detection for Decoding-time Jailbreak Defense cites this paper.

TrajGuard: Streaming Hidden-state Trajectory Detection for Decoding-time Jailbreak Defense Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:50.778962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T18:26:19.922383Z digest=sha256:0592767c37d9d44dbfd0ba8de68afba7e71e82a81dbd4064d0d6b4475125129f

Observation 3d58642f-7d0f-43f0-9c67-768094314bbb · inbound

GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking cites this paper.

GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:45:37.340698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:40:59.993298Z digest=sha256:75b1af06b0ebb2b395ea634d9a6af0381cc93d876eba34328ae6771dde428c5b

Observation 25c6cce4-a38e-4317-b3ab-66456ab89088 · inbound

TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs cites this paper.

TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:26:00.015367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:02:52.006859Z digest=sha256:c9176e37cc7f20dba860b2ccd5354e6d9b92662d5f140ad798a9b07ef268b9c2

Observation 7b10e68e-4c58-4538-949f-52ec49ce67bb · inbound

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff cites this paper.

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-12T20:09:57.922788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:09:57.922788Z digest=sha256:397440c593ea85814fcbf4954e41da46865ed3c21739784dea1f8d44b8870a37

Observation f2448a1f-80b8-4bbe-b1be-faeaf618d404 · inbound

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection cites this paper.

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:35:18.931019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T11:32:10.126062Z digest=sha256:b04db64204d93276a2cf52ea2d32ca6d8ee82855801f14cfa93c208abc4015ce

Observation 75fea871-0b2c-455e-b68a-69d3a924bc36 · inbound

A Systematic Study of Training-Free Methods for Trustworthy Large Language Models cites this paper.

A Systematic Study of Training-Free Methods for Trustworthy Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:22:37.353269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T08:19:42.671690Z digest=sha256:ad2c2ffe5aa02b49d3680331365919716be85e96a337f424679d8ccd121ca287

Observation c0804e6e-5974-4d59-9fc7-fcc348e25e30 · inbound

Jailbreaking Large Language Models with Morality Attacks cites this paper.

Jailbreaking Large Language Models with Morality Attacks Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:21:27.030426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T06:17:14.242224Z digest=sha256:2bbd9d420ae515ad17b90fc3f9633bd040ef0ee585e9e0be34691fcddb7a194f

Observation 654cd810-3283-4979-be35-e00290ea9fe6 · inbound

SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models cites this paper.

SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:06:04.411354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T02:21:29.463149Z digest=sha256:44ab31d699cce8601674469f886329219637ae76533ac50162d5031836510530

Observation c698fb39-f359-4ccf-bc76-66c16fef678f · inbound

Automation-Exploit: A Multi-Agent LLM Framework for Adaptive Offensive Security with Digital Twin-Based Risk-Mitigated Exploitation cites this paper.

Automation-Exploit: A Multi-Agent LLM Framework for Adaptive Offensive Security with Digital Twin-Based Risk-Mitigated Exploitation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:09.476633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T11:45:21.608035Z digest=sha256:059b0bb7a4f485eb931a136af647e9eb5296bef0629fbf5cc2166a043029a3bd

Observation 63e4218e-7582-4fa1-9f5b-30946ea95025 · inbound

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework cites this paper.

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:51:09.567702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T07:53:13.746141Z digest=sha256:559714994826f62e5c04cf27f305f4357ea9315112a6b07c0b4c6151d4e81cf8

Observation 5d7afb9e-d275-447f-9f23-c01aedbff02b · inbound

Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models cites this paper.

Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:01:09.629032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-09T20:39:39.225898Z digest=sha256:7999967cb6866d7e86b6cb0dc25c04c02fedcc558f53af186896c1a31375d7e8

Observation 598a05e0-a41c-42a9-858b-69653353566e · inbound

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms cites this paper.

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:51:43.987363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T01:46:49.586630Z digest=sha256:dfa0e1905c524cec644220bfb85c2e67388ff7384ac8fce85f83f3a16399870a

Observation 833c9c08-b805-4c2d-828c-8bd4f095fd6c · inbound

PQR: A Framework to Generate Diverse and Realistic User Queries that Elicit QA Agent Failures cites this paper.

PQR: A Framework to Generate Diverse and Realistic User Queries that Elicit QA Agent Failures Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:03:36.763330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-20T18:03:07.646917Z digest=sha256:374f6f636ad16ce6dcd608abc22a9d961e5b77215a4d161970ef9f272b460d39

Observation c766b4e8-3c80-4acc-acc4-12cc8b8205e3 · inbound

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models cites this paper.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.699039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:118426bc496459dd3a4126fc739ce66e6f97f419a84f5911dd8b1e78363b8726

Observation 440422c9-2f00-4359-a058-80339a4cc457 · inbound

THRD: A Training-Free Multi-Turn Defense Framework for Jailbreak Attacks on Large Language Models cites this paper.

THRD: A Training-Free Multi-Turn Defense Framework for Jailbreak Attacks on Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:19.173126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T15:05:31.392966Z digest=sha256:a9d70e425f8d97b3fbd79368824d57399c767dcc27a4be43de34654420538f60

Observation 0806e1bd-bc88-4890-a644-04a84df5052a · inbound

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks cites this paper.

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:26:59.309811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T01:16:07.252429Z digest=sha256:55a1a00388cba149f93c96aa847a80b153f9de6efbf75fbc763745b9ed64f1e6

Observation 7ad5939a-73f8-40b2-993c-5fe2b2a0181f · inbound

A Layered Security Framework Against Prompt Injection in RAG-Based Chatbots cites this paper.

A Layered Security Framework Against Prompt Injection in RAG-Based Chatbots Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T02:09:22.521177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T20:00:14.036515Z digest=sha256:6d33b621d8438bd431e6a9944c5a5cd0696eec4935ab18fd87607a477c5db841

Observation 89595de1-b9ad-4080-a147-1ed5fd276557 · inbound

Investigating The Security of Modern AI and Cloud Infrastructure cites this paper.

Investigating The Security of Modern AI and Cloud Infrastructure Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 137

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:29:42.325046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T11:31:39.910784Z digest=sha256:4e59c090777a7af64e73828be384eaeb4e906f156184de658e1f38453dd90359

Observation e938f8b3-7d40-4f1f-9857-d41e30c67f32 · inbound

ToxiREX: A Dataset on Toxic REasoning in ConteXt cites this paper.

ToxiREX: A Dataset on Toxic REasoning in ConteXt Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 223

Resolution
verified exact
arxiv_id, observed 2026-06-29T04:43:07.121701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-29T04:33:18.794505Z digest=sha256:d0dd7db7a930c5a774c6251c187704ab6d4abb6b8100256459715da35bd7230e

Observation ccf75546-7565-452c-b4bb-2286973270ba · inbound

Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety cites this paper.

Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 104

Resolution
unresolved
no resolver link, observed 2026-07-12T14:28:50.627444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T14:28:50.627444Z digest=sha256:efe570b1963b87e22e89defbee402d60b305f22338d4add8f1bb0c3e72962f53

Observation 121b3da8-53a6-49bf-a5ae-cf67e5e61b74 · inbound

Mitigating Taint-Style Vulnerabilities in MCP Servers via Security-Aware Tool Descriptions cites this paper.

Mitigating Taint-Style Vulnerabilities in MCP Servers via Security-Aware Tool Descriptions Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-09T10:26:11.097166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-09T10:22:23.782469Z digest=sha256:95e77d4dfd24c16ceeb704803156a0484806580f76b2dab48dc83272d4b49d63

Observation dde8f715-429a-4a7b-a486-671e594e1b93 · inbound

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions cites this paper.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:01.084617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:01.084617Z digest=sha256:d560c1c8685609384ee0578c77705ad0cc3734ee8ebdc1bfda4970291c7b829c

Observation 7a179e9a-bb3b-445a-b683-f7a1f9e23786 · inbound

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs cites this paper.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:05.960205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:05.960205Z digest=sha256:00deb527849d523f4d616ec523677ff8b1cbce01f8143bbca1dae806cd5775ad