Pith. sign in

Paper Citation Record · LEDGER

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures

As of 9 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2506.07402.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07402 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:40:20.752585Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:55:03.319920Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7ac0cda6-c77c-48c6-99e7-985e17627304 · outbound

This paper cites Universal language model fine-tuning for text classifica- tion, 2018.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Universal language model fine-tuning for text classifica- tion, 2018

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:21.505766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:40:20.514067Z digest=sha256:9dd0d5944960de484f0520a03af3e26573c70ad54a60c526b494ac337ba4bd38

Observation 7a8c6123-ac07-45b4-aaef-ea22fed4a683 · outbound

This paper cites Training language models to follow instructions with human feedback.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Training language models to follow instructions with human feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.519392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.519392Z digest=sha256:2e888343a11c51154f375732af236ed2fb7ba3152dbf91310820a9c877f01c30

Observation d7e06a19-25a4-49ea-9cf4-46033b91e68d · outbound

This paper cites Scaling instruction-finetuned language models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Scaling instruction-finetuned language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.525164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.525164Z digest=sha256:25312dd187e606f6121248715b46c6b538883fe50ae7c5b54ca438afa4d2cbe3

Observation f687c9cc-53af-4811-acf1-5026fc9bf588 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Fine-Tuning Language Models from Human Preferences

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.529978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.529978Z digest=sha256:472e28110009862b39b8f0e8c0920e4d90496b9da1f29c8c2bd95cb651e752e6

Observation 6938c8e7-dd8a-473b-a01a-e7b0717c189e · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Direct preference optimization: Your language model is secretly a reward model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.535236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.535236Z digest=sha256:b612abc043ca6693cab0428b5dff493cf12eed2207fa07f14064da6c654f1905

Observation be8403d9-f7ee-4916-a14e-cf1608a7fa44 · outbound

This paper cites Jailbreakchat.com, 2023.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Jailbreakchat.com, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:21.439636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:40:20.539989Z digest=sha256:de3e8f07503429e04422164446cd43a74d6dd37fef650abe32fc8e4bed3a0da0

Observation f6ff2025-0476-4e55-9c57-d49f5a91cb13 · outbound

This paper cites Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36:80079–80110, 2023.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36:80079–80110, 2023

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.545428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.545428Z digest=sha256:a0aa8359a829f37a4816707ae679ef9ddc22666ed36fcf2049a0bfaf19b2246b

Observation 2f1a9f34-036d-4335-8a32-9f3ab8954e01 · outbound

This paper cites Are aligned neural networks adversarially aligned? Advances in Neural Information Processing Systems, 36:61478– 61500, 2023.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Are aligned neural networks adversarially aligned? Advances in Neural Information Processing Systems, 36:61478– 61500, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:21.415500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:40:20.550874Z digest=sha256:c18d4e111e316109234006838baa4232b13c2d2fe6222f1d89b656baae4e29c0

Observation bd3b9ee6-56ab-4793-ba5a-dc13d7fd6fcf · outbound

This paper cites do anything now.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures do anything now

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.555354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.555354Z digest=sha256:91b0db8ea386eddc7c30521ea5f98b2dd01b9fc378f558c0867a83dabdd1dbdc

Observation a1753045-f0d9-442a-aecf-64ed9abd11ff · outbound

This paper cites Don’t listen to me: understanding and exploring jailbreak prompts of large language models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Don’t listen to me: understanding and exploring jailbreak prompts of large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:21.390445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:40:20.560611Z digest=sha256:9264e6538816eebcbea693590ed179bad86a27cb7c1290e9eb87f78ad3827bc9

Observation ac5bca8e-c73d-4f9e-90ca-fa8bed0495f8 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.565977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.565977Z digest=sha256:d0a63181d04a6cb6d84d74e94668748f29474429838ef8f48b1d2bd9183d262a

Observation 36c6d1dc-e66c-4e72-868b-e30305efdd7a · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.570649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.570649Z digest=sha256:ca391d566c60a46325ee71571bf87cf288b20f07afbce00532dfb452e324ca74

Observation 7d463ea2-dc58-4f1d-b807-d3b19f063b03 · outbound

This paper cites Tree of attacks: Jailbreaking black-box llms automatically.Advances in Neural Information Processing Systems, 37:61065–61105, 2024.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Tree of attacks: Jailbreaking black-box llms automatically.Advances in Neural Information Processing Systems, 37:61065–61105, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.575426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.575426Z digest=sha256:1c7969248b94e1f5ad8ee2d239f7411756aad30d47fe8dc326499852e119b5f2

Observation ff06be60-b1ab-4ac7-b4c7-f3d678f2af6b · outbound

This paper cites MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.580018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.580018Z digest=sha256:0aaa7a4b073e92da964291ced9e669da3e1b9d6d7bf7a61cdb8e9e8be64ef478

Observation afa7ba33-3fe7-41de-877d-2461d41a0ce3 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.585362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.585362Z digest=sha256:42ddcd463375b5103e5c3e04245037d660c0d2d599a35b04ca7ed5143b54510f

Observation d202259b-976c-48e2-8129-af12e45802bc · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.590342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.590342Z digest=sha256:9b469f5cc34e4c1e323bea33489e5de1d4952311ce2713ae5cdca2656a7014a1

Observation da3d46d1-a7b7-4829-a1a1-a80d0775e3ca · outbound

This paper cites AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.595942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.595942Z digest=sha256:68a4f724f112c1e2f854b1aff5348bcabe6cfa55956d92d1c4ee4a34da5eb14e

Observation 8cd526e5-5d8d-4e82-91ab-2b9b4db91aa4 · outbound

This paper cites Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.601664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.601664Z digest=sha256:930bbc335ad5393b93408cdfbe5b799fc73ee09f156e631e441064374db9d67e

Observation 595452c4-6357-4760-8f52-30d432d8a57b · outbound

This paper cites Weak-to-Strong Jailbreaking on Large Language Models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Weak-to-Strong Jailbreaking on Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.606600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.606600Z digest=sha256:af24c79450522727c6969d2387b2c31ca7f9ada9bced66ff5a70ea3189d30567

Observation 81904555-64f5-41c7-a0fe-bfecc3e67c03 · outbound

This paper cites AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.611175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.611175Z digest=sha256:0e26a73022294625abe0f7ac6531c37c033042698f915d2ecb1b85b1db56a87f

Observation 72aa392a-d53a-4eb3-8989-88f9158796ad · outbound

This paper cites Improved Techniques for Optimization-Based Jailbreaking on Large Language Models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Improved Techniques for Optimization-Based Jailbreaking on Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.616360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.616360Z digest=sha256:e351b7ed0f535fd68a81669bc6279fb7714ce6996598cea812382530cb084c5f

Observation 72ba783a-5b5a-4037-a109-1035b69cdfe8 · outbound

This paper cites Don't Say No: Jailbreaking LLM by Suppressing Refusal.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Don't Say No: Jailbreaking LLM by Suppressing Refusal

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.621417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.621417Z digest=sha256:779ca27fd15ebb925f304d2ff0627cc3db0bf9fb1e4b75554dc8e55876b1a096

Observation 95cbf8a9-6f0a-4ea8-b67d-d55cec96b58f · outbound

This paper cites AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.626685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.626685Z digest=sha256:27730bd255387c7b8b12c6877b33a06bf50991ad36b668a7b0189f6a3bcd029e

Observation 37d45666-6cd8-4238-a472-ce1bc8adba7e · outbound

This paper cites AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.632046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.632046Z digest=sha256:fa5c7f6833f771a74d9a52bb9345b7050e07e7135f25f303db51a30ad4af17ae

Observation 7d542fe5-5475-47de-b4ff-013c8b89ec90 · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.636809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.636809Z digest=sha256:37545478555aa9d67136a8adec808f178b11cbb94d37353efb37e3ee08371df7

Observation 9bcc24c3-5661-475e-b745-7e0a568132f8 · outbound

This paper cites Llm jailbreak attack versus defense techniques–a comprehensive study.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Llm jailbreak attack versus defense techniques–a comprehensive study

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:21.365440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:40:20.642041Z digest=sha256:1e6194ef144b9986ba8ea27abc21c969492fe308c3a56b955739d2a81a44f58f

Observation ec8ff256-1793-4bab-b989-42bd5d49e3b0 · outbound

This paper cites LLM-Safety Evaluations Lack Robustness.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures LLM-Safety Evaluations Lack Robustness

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.646900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.646900Z digest=sha256:1a200b0263648e040904298648021373cd4ffac3f51ff22d2ad27ff1b1452881

Observation b2eaf344-2622-4f66-b217-4850c0eb5904 · outbound

This paper cites JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.651700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.651700Z digest=sha256:eeed345ba11c61ce9d130829b53aa8e31a53d1444fd1f96fe70e653d4ee31fcb

Observation f20d886e-3a37-4af6-864a-5e093bcc8e5e · outbound

This paper cites Sg-bench: Evaluating llm safety generalization across diverse tasks and prompt types.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Sg-bench: Evaluating llm safety generalization across diverse tasks and prompt types

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:21.349668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:40:20.657041Z digest=sha256:a0623c8c5f9449fabf8b97f4a6ccfe428a78a01e9f19377d26c62aef5f36e02f

Observation ebeab4e4-5f8f-4310-a1ba-9de7ff330d25 · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.661675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.661675Z digest=sha256:17c179581c111ac727e3b2e862cc248d891a3a1dae636164531efb43b3844dba

Observation 51c51fb7-d940-44ea-83c4-8445e1beb86d · outbound

This paper cites JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.666525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.666525Z digest=sha256:b7cf071aec1d740787ab7a2b2e951b8e90dfbfbebd6384b34b399b2c0ae385c7

Observation 3f30fa75-a85a-42b8-b79d-0bb40852e600 · outbound

This paper cites Clas 2024: The competition for llm and agent safety.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Clas 2024: The competition for llm and agent safety

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:21.333247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:40:20.671409Z digest=sha256:f4b5d935ee86e7e92e272e189767b06037899f95661e5dd3ec4077d8ce92d358

Observation 7db82a51-965f-4deb-8037-686fccefef3d · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures A StrongREJECT for Empty Jailbreaks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.676687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.676687Z digest=sha256:c4261523036fa52960243df77d9b136aa0bda07c73471bc3874a62db5ca67d9e

Observation 606f5adb-80f0-41a5-8703-0e69f61b6ff6 · outbound

This paper cites The art of saying no: Contextual noncompliance in language models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures The art of saying no: Contextual noncompliance in language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.682110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.682110Z digest=sha256:49a973d92e6567f904a0437bf2c2cc8685a80a381d425673f40d6587cb4db372

Observation 931ffb9d-d955-4d82-b723-9b4705eebfc9 · outbound

This paper cites How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:21.307575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:40:20.686767Z digest=sha256:368e0254cce024f5072aebe625ca5e44c1ddc5a5cf93e81e0305af0f77f21431

Observation a0dc3ecf-570a-4d1e-ae01-7785e1e52f36 · outbound

This paper cites Jailbreaking large language models with symbolic mathematics, 2024.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Jailbreaking large language models with symbolic mathematics, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:21.291327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:40:20.691795Z digest=sha256:bf66f2fb4d613796e759a622b1f045f35a90737e43e996c1fbbb08e140093922

Observation 5b995f12-018a-4ea8-83a6-6fbb92a7a99d · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.696608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.696608Z digest=sha256:5168370550187c3862429242d3a3a19f2e54d7aeb320acf1b7ad77b3086ff7ea

Observation 34268b1e-ffbf-4159-9ee4-6aeecca7a4bf · outbound

This paper cites Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.701570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.701570Z digest=sha256:29ed5f535af43ea7c64f8ea8bcee9831bd321758b71e160ac22c4e73a399aed9

Observation 37041393-4ed4-4b92-a263-8c859f7bd858 · outbound

This paper cites Rethinking How to Evaluate Language Model Jailbreak.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Rethinking How to Evaluate Language Model Jailbreak

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:40:20.926237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:40:20.706822Z digest=sha256:3469136de82acc6b6b3150b3912d770e8ee78023e9e3f5ec944f92ad456a50d6

Observation 2a7751e3-04d6-4088-b57d-1229ae8704ae · outbound

This paper cites SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.711744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.711744Z digest=sha256:cbe54cf80b2a0425ecde81eff0cccb38db4527fa23e8b516cddf1d791aaed201

Observation a1070a30-5b1e-4ab6-992e-ec2f7ad6ff83 · outbound

This paper cites The Jailbreak Tax: How Useful are Your Jailbreak Outputs?.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures The Jailbreak Tax: How Useful are Your Jailbreak Outputs?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.716732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.716732Z digest=sha256:793ca0910fd0d59dcc5f2cff9cb1eebc616709576b64874753a47270cd8e0ff8

Observation 26835b23-1641-4939-ab42-09bdb134a4a5 · outbound

This paper cites AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.722684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.722684Z digest=sha256:1d47d9e44655da23efef2e5234ee5db2b114e525555b4d15922b3ff507890355

Observation 2912aaeb-2854-4b65-8fa6-47bfc05b8e3b · outbound

This paper cites Multimodal Situational Safety.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Multimodal Situational Safety

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.727309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.727309Z digest=sha256:a34eb60264472cd7a7d868819d360b92be5af4522aa3a93e689ac04fc8d52206

Observation b139af91-eb49-41c9-8e53-aad5732b7f87 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.732373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.732373Z digest=sha256:f6e5828d443e6858552e6984bcc18e6abac92a7396a820007d994dd0d5ee2fb6

Observation 19b5f8c9-896d-4480-b287-6ebff5ed5fad · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.737235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.737235Z digest=sha256:c0738a2765b0e2532a57e3b14a65f81bc33ab1cce5f3401bb668945c6bd31d0a

Observation d8e565ca-a31a-48a8-acc5-b88382c9daf0 · outbound

This paper cites Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.742100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.742100Z digest=sha256:9e385abdf2715fc4f283d32ddc5933ee7a883cc0697389905c172ec53ceda014

Observation f9ee1f79-47cb-4735-aa43-bfa24bca4413 · outbound

This paper cites Multilingual Jailbreak Challenges in Large Language Models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Multilingual Jailbreak Challenges in Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.747279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.747279Z digest=sha256:865934e0a4f2c5d873321c434aa4a00215a7825fbed86f943a7d626cf889c3d3

Observation c549df12-5d4b-49ff-9913-808afde46e72 · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.752585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.752585Z digest=sha256:2fbc2b69107224e5c74eac0038132c4bb59e58f4fb31147860e726cb9dfbd2a8

Pith citing papers

Observation 394b7429-0807-402b-917e-b5967f811225 · inbound

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions cites this paper.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures

Reference 131

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:03.319920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:03.319920Z digest=sha256:adfb1807f62b3d48d8c299f0c925214e997b2f28eacfb356b03095c14e713189