Pith. sign in

Paper Citation Record · LEDGER

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions

As of 17 August 2026, this Paper Citation Record lists 100 of 138 outbound references and 0 inbound Pith citation observations for arXiv:2607.17152.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.17152 v1

Coverage vector

measured 100 of 138 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:55:01.084617Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 138 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 18002605-9fa5-4bad-a6f5-10bafd6ca0ca · outbound

This paper cites Scaling Learning Algorithms Towards.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Scaling Learning Algorithms Towards

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:52.248609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:52.248609Z digest=sha256:714257978902be571beca408d6031c0c9fcabab6e8ddbeaedc815edae235da50

Observation 1ff65e5b-8ed1-499c-bba7-ac446f71310f · outbound

This paper cites and Osindero, Simon and Teh, Yee Whye , journal =.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions and Osindero, Simon and Teh, Yee Whye , journal =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:52.365173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:52.365173Z digest=sha256:aae1e848e13064716403a4fc2b907dca15e4d77f9a9ce6a4a8eeab3a38f5524b

Observation 848d4d36-00b6-42d8-8c6c-fee4eb003cb9 · outbound

This paper cites 2016 , publisher=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2016 , publisher=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:52.550470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:52.550470Z digest=sha256:8a60e89812a63fcb617da154cb854445e8740245e7521bcca2a414e48338f0c8

Observation 3e822bc7-0429-404b-8301-fa5f5662dda7 · outbound

This paper cites ACM computing surveys , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions ACM computing surveys , volume=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:52.671464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:52.671464Z digest=sha256:8f6cae06b916b2431061c3c47dbd8136da37c0e814be4a54d3d0a390050fd87a

Observation e1afd918-d576-4d97-b909-ecaba04672d1 · outbound

This paper cites LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:52.804286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:52.804286Z digest=sha256:210558c9f378b10b44c817ea0603b505fc0b21d0f6330b4aae0200cb508f89d0

Observation 7c3bdb28-b686-421f-ac4d-5bb66a412136 · outbound

This paper cites AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:52.928929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:52.928929Z digest=sha256:87f93df6dc90bdd64837b5efa22e3d5e03a2ace4b920d66f1a677c541fe8e3ac

Observation 89767f0f-59e4-4f9b-8fc7-a2ffc0dc38e9 · outbound

This paper cites Multimodal Situational Safety.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Multimodal Situational Safety

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:53.054207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:53.054207Z digest=sha256:b1f25c2618d0c42b15ff6b93156466946c3b53f8208179092669cc277320d815

Observation 967d7966-324c-44d1-abd9-4920a7cb276b · outbound

This paper cites The Jailbreak Tax: How Useful are Your Jailbreak Outputs?.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions The Jailbreak Tax: How Useful are Your Jailbreak Outputs?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:53.171272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:53.171272Z digest=sha256:792842ba7dac0c01136b5623e1f551d1e69968b3ede1552297ccf62337507652

Observation 7104ac5d-abc2-4143-99b7-f91bcc176f93 · outbound

This paper cites "Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jailbreak.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions "Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jailbreak

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:53.295354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:53.295354Z digest=sha256:696937ad86a512ad0d9d61ef8d17b80aa3c0ae6e2dc753024d29e111528e942a

Observation 12eeab7a-76cb-43ea-810f-b49b34b23d8d · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Advances in Neural Information Processing Systems , volume=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:53.455230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:53.455230Z digest=sha256:3769a1b49c8cb8827d1f241c63e13566d40e6147f14b82102740552117e2810d

Observation 8971f3d3-dace-45e6-a8cd-216a155e3ebb · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Advances in Neural Information Processing Systems , volume=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:53.564076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:53.564076Z digest=sha256:1afa5c8e91f526bece9ea0523448cf40a73505456596f61b176fb56d266de24c

Observation 963f3aaf-ca90-4070-af10-6a0f2444964d · outbound

This paper cites do anything now.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions do anything now

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:53.703136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:53.703136Z digest=sha256:89dff82f9423be32f748732dcbfd482ffef21e880435c80c73f3a94034a66ae1

Observation d18954ac-a5c4-445f-925e-04e051baeb71 · outbound

This paper cites 33rd USENIX Security Symposium (USENIX Security 24) , pages=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 33rd USENIX Security Symposium (USENIX Security 24) , pages=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:53.827434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:53.827434Z digest=sha256:45dca05b200b98020db213b817e412e642837a09bd6678b29129b84240e5051f

Observation 2daa9fdb-a5dd-411a-b9b3-680399d26b00 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:53.960533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:53.960533Z digest=sha256:1a5fe83dc36e01e9464f4fa1fe8a7efb3266ee6cf355e8b2265566b4a48a51c6

Observation e7abd134-6d0c-436a-8bca-6f55970e726f · outbound

This paper cites 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) , pages=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) , pages=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:54.084873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:54.084873Z digest=sha256:58b95789851e0709e8cd6540b64b156672b09615a8467d3bb820aec8dfbddcad

Observation 74e2fb86-80b9-4124-ac9c-322e00e414e4 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Advances in Neural Information Processing Systems , volume=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:54.215365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:54.215365Z digest=sha256:88a66d36ffe826ec59ed273e7be0f55535b17c0c802dd98bba12dee1e3efa08a

Observation 8cc89bc9-ae29-4864-8b67-7c4d443dfd22 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:54.352289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:54.352289Z digest=sha256:640fab08161faa4181dbf66256c834d1b889891af4f4347243277b8973da8cba

Observation fe62a189-eca0-4386-96c1-2a54cc8a66c3 · outbound

This paper cites MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:54.483085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:54.483085Z digest=sha256:7d3d0c59a0aeba45366e8dbf7fa44e1805f8e3b2a397c388249be6d46aa84ac5

Observation a516a966-e485-483f-a54b-8edf16232ae1 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:54.604317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:54.604317Z digest=sha256:9be1e7b935ba269629577776fda8688408c12c7c9a660b7eba808272e778877c

Observation f761a488-706d-41aa-b7b8-a32b2dff86de · outbound

This paper cites an unresolved cited work.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:54.734577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:54.734577Z digest=sha256:bb2ea96e1bf254c2326ebc21564fe41ee04b3023c0d7ab8ee505c1c32c8b601b

Observation bb9a97c1-7f8f-4af3-b18f-f5ae0787f806 · outbound

This paper cites Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:54.885903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:54.885903Z digest=sha256:c8e1b2e66f57d04834b52c9c766edc3d16b1b44d6054f8842f64507e894eb09b

Observation 52fadcb1-083a-450b-ba00-b19f4d4df369 · outbound

This paper cites Weak-to-Strong Jailbreaking on Large Language Models.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Weak-to-Strong Jailbreaking on Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:55.053927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:55.053927Z digest=sha256:73d6e67a19d262a72e426fd160d2082b582ef3a26ab6f42289a2cabafcb42a1e

Observation 634f7881-0dcc-477c-b3cc-0de8ad8bb93e · outbound

This paper cites AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:55.186195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:55.186195Z digest=sha256:9243b46e949c1e9f816df2201a3f517553c083443e36d00c10d33ded51a91ee4

Observation ee5f4810-1ec9-4d50-b3aa-eebb2220b798 · outbound

This paper cites Improved Techniques for Optimization-Based Jailbreaking on Large Language Models.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Improved Techniques for Optimization-Based Jailbreaking on Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:55.306378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:55.306378Z digest=sha256:89a3433324f4a809112ad8986163360026d5b756b420ce13f43d645b1d39bbdc

Observation bac4df12-24e3-4676-83c9-33794bd5fab2 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2025 , pages=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Findings of the Association for Computational Linguistics: ACL 2025 , pages=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:55.442444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:55.442444Z digest=sha256:84b7aee2dee128e13722cf9889461279932e5e45df708dc96a9b2aa8df7a5fbb

Observation 58d8a426-7b61-4a49-b511-3bdb7575b589 · outbound

This paper cites an unresolved cited work.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:55.544716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:55.544716Z digest=sha256:5afc9fecbab8a985238783439d24e59bd172ce9cb9b2fdaa2c2d0364b3ec8872

Observation 2cf06490-76ad-4784-9219-b3c6069fff03 · outbound

This paper cites Edward Suh and Yevgeniy Vorobeychik and Zhuoqing Mao and Somesh Jha and Patrick McDaniel and Huan Sun and Bo Li and Chaowei Xiao , booktitle=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Edward Suh and Yevgeniy Vorobeychik and Zhuoqing Mao and Somesh Jha and Patrick McDaniel and Huan Sun and Bo Li and Chaowei Xiao , booktitle=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:55.610459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:55.610459Z digest=sha256:4a21bad08a31f75be7271028d101ad081e54a211065d2d66354ddbb221d45cb8

Observation 3cdea44b-aaf9-4245-82ce-4fc125db9ddf · outbound

This paper cites 2018 , eprint=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2018 , eprint=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:55.678718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:55.678718Z digest=sha256:933b820cad6abc28fa4ef23e2b3a84047b26397c627a4d5f73ac7c28beae0b20

Observation aac44d2d-62ff-4b5a-b33a-0efd711ba92e · outbound

This paper cites Journal of Machine Learning Research , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Journal of Machine Learning Research , volume=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:55.797068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:55.797068Z digest=sha256:a111f9385d1144d9dcc6b58cf5401f668787f145e8bf9b4669decfaf7ae6bd27

Observation 138d8da7-079f-4616-af3e-0e4f7f0f4f31 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Advances in Neural Information Processing Systems , volume=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:55.870808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:55.870808Z digest=sha256:72d055e6dd22cbeaeb80bd24cb0b2dc77b527d078805709c63bdd7dd9bfd463a

Observation 8ba6f0e7-97e4-4c31-932b-baaf6070ab1f · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Fine-Tuning Language Models from Human Preferences

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:55.935710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:55.935710Z digest=sha256:ac1b461f669dd9cf0dc738a35a94531c385c736538e692df2cba1dc8fbcd2200

Observation 24f8ae2d-9fce-46cc-b4ab-41964e720e21 · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:56.006123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:56.006123Z digest=sha256:c21771e821e2a236b17903783e6ee49cb3808542f395e698974ab301b555110b

Observation 33f297fe-3784-4bbc-abb0-e779cb68245c · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:56.071336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:56.071336Z digest=sha256:7a9b60dec22dc72d85d600094994e427b8640cc2880be08cc8dfac11a33d0505

Observation 6bcb71d8-d5bc-434f-abc4-4fea114ae1a6 · outbound

This paper cites arXiv e-prints , pages=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions arXiv e-prints , pages=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:56.181744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:56.181744Z digest=sha256:fe701e68d9680c4c61b25b0f0e01bf83b1180d9995f25f0d130382d384084acd

Observation e4bc5f1f-d17b-4223-864c-af779694ee32 · outbound

This paper cites LLM-Safety Evaluations Lack Robustness.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions LLM-Safety Evaluations Lack Robustness

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:56.256964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:56.256964Z digest=sha256:54b389edaa2cf012fcc2cde92aa1221f0890d08d87805e5d4c8f58ede0792b6f

Observation 7870a03d-dc66-4da8-8a26-13a488e66d5c · outbound

This paper cites A Survey of Hallucination in Large Foundation Models.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions A Survey of Hallucination in Large Foundation Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:56.305693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:56.305693Z digest=sha256:38abce1aecad4d1dee1a69fd878ebe43495c3f3986f33ef5b0287bc8c626e44b

Observation ea86d252-d43c-4c70-bd5d-ca91156f2953 · outbound

This paper cites ACM Transactions on Information Systems , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions ACM Transactions on Information Systems , volume=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:56.364474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:56.364474Z digest=sha256:0e0acc7847d797d2c139bd240bed2eb026c4ecc7c10624048ce989e962f0f325

Observation 7ae36b45-8dd6-4030-b929-9faba43241e6 · outbound

This paper cites Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:56.415531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:56.415531Z digest=sha256:e959b08ab6b90df7b9fb7eefdf92e9fd786cc757b530cf89950b6c2f5a99d0ca

Observation c1f2901e-acff-42b6-b9bb-c522c5a80d7a · outbound

This paper cites Advances in neural information processing systems , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Advances in neural information processing systems , volume=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:56.514885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:56.514885Z digest=sha256:180f99beedb38296e7e86fc35cc9b91e20e1a545bf1f0a55bef6a2180767c829

Observation b430ccb1-6a13-4671-825e-1544dd768caa · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:56.575779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:56.575779Z digest=sha256:a1caadf839e04ddc2e556278a2fa4f676dc0c93027df99de9204713bd634f152

Observation 4c7e6fec-e06b-48d0-92d1-2ce08b9f29de · outbound

This paper cites JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:56.625375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:56.625375Z digest=sha256:8f7fcbef5748c9301eb225d37f40d8fac7aef69df7ce80610a447cf1b5a09d9a

Observation 0f559ffc-5998-47bf-9b49-9a3f6b9aec65 · outbound

This paper cites NeurIPS 2024 Competition Track , year=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions NeurIPS 2024 Competition Track , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:56.682990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:56.682990Z digest=sha256:034f8fa98d2ebae47287ed28a1bfee0186553cbeb75b763da1d8de97ac3c12c4

Observation cf882857-a8cc-4cca-82a7-b80739ec12f0 · outbound

This paper cites Rethinking How to Evaluate Language Model Jailbreak.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Rethinking How to Evaluate Language Model Jailbreak

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:56.732635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:56.732635Z digest=sha256:ec39f936f7c3e1d848e34ccc0ed82f077dd1e3d02c576d733b1d726ced580f4c

Observation 03044fc8-b0f7-4f87-8616-0f7d5c19e5bf · outbound

This paper cites 2023 , url =.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2023 , url =

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:56.829238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:56.829238Z digest=sha256:ed22edb6e1d15335207643f12e6dda6160856c9474620f830f3c61b664e40ce1

Observation d4e22f94-39ae-4794-9235-9c4b44865aee · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:56.917242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:56.917242Z digest=sha256:f8bcd20f7ba5d2c486d8fc2f6a2ec693ae2bd76410a1f4c5b221e5ee7a563972

Observation e6e2d029-4e1d-4555-909a-f4d2750d635e · outbound

This paper cites 2024 , eprint=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2024 , eprint=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:57.009764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:57.009764Z digest=sha256:485b6352a7af7eef233a4cee58fa0de73b9c94ff6d6682eb466c77ca2a607657

Observation bb7d12f4-52e4-4dce-8a41-62bcee50f814 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:57.057116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:57.057116Z digest=sha256:662bded2b0558e2ad8bbab79f8d3056f6ae815ff5bba36af8c27233e3624708e

Observation 3406f3d7-d8c9-4e94-8c11-f3745de7deff · outbound

This paper cites Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:57.117807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:57.117807Z digest=sha256:4db0e085fee6a2834d46e032375351d367dbf39ddffb7bb7c7aa7c9ec3a0709c

Observation 0d12b4db-f03f-48e6-83eb-facc8d89715b · outbound

This paper cites SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:57.199932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:57.199932Z digest=sha256:236496a524cceb16848ea9a2ea7551b37a02f62746e69085cbe0f28daab4b77f

Observation a1bb377c-d590-423c-bf0e-15187944036b · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions A StrongREJECT for Empty Jailbreaks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:57.280117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:57.280117Z digest=sha256:f2429ccd190179a292c29424fe5eb6c554de93a47b12b8d199ff09edaa358488

Observation 7742f2ee-ca4d-4638-82e5-45b309e1d408 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Advances in Neural Information Processing Systems , volume=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:57.338164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:57.338164Z digest=sha256:2e218f402aade252325c5fa772075ddf352ee4cdf0593b6c6124eab515afcb84

Observation adf2f54f-3431-4866-8881-f12569d34608 · outbound

This paper cites The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:57.407015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:57.407015Z digest=sha256:fb13dfa68aa7301f997e14f06a60299a94d23a0a068564732b0c9d7be7834c24

Observation ab5d1d69-c259-4ad9-8662-d6e660ffc989 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:57.490593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:57.490593Z digest=sha256:bf7bf06c11c2ff9a740d9ebbe5b50741cf6f54b6e0f94979126f96999f9e3560

Observation 69ef5eac-fa14-4423-82d8-da73b3f6e575 · outbound

This paper cites Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:57.589967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:57.589967Z digest=sha256:ac9c072dd8849831ff312b3aca5f6d5767a10079fd9595007b242387fe41728b

Observation 4088aabc-2307-467f-8a8a-d350cbac629f · outbound

This paper cites Multilingual Jailbreak Challenges in Large Language Models.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Multilingual Jailbreak Challenges in Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:57.637430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:57.637430Z digest=sha256:77d2a54bc81c4876aa85941a97ba95bffc3758938418b372d3df03353680484d

Observation 754b9883-c0c1-4889-9df2-8d64b1689f7a · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Advances in Neural Information Processing Systems , volume=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:57.720048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:57.720048Z digest=sha256:8a4894f91289ffb36ad028e781bd77a8b56922c2279cb7df050ff5365c873115

Observation e94f6a3c-d03e-4e58-8f17-d7a4b741ab86 · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:57.776258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:57.776258Z digest=sha256:4409184189011a66d9f95f9de1dc25206809c497555cb9876f8fc2c79bb5443a

Observation 5650bd3f-f3ef-4e01-9f12-a203a04b2d32 · outbound

This paper cites ACM Transactions on Knowledge Discovery from Data , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions ACM Transactions on Knowledge Discovery from Data , volume=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:57.831509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:57.831509Z digest=sha256:e843f45b0b12a2dc703b03eba2449aa92701a255380e0d496249eb697973deca

Observation b2aaa0d2-6ef4-4a4a-a0ec-07d9dfee04ad · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:57.913572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:57.913572Z digest=sha256:e6d675c69330a69d57e9488dd3ec5bbaaefba60c2ba46f1f5cb935b199186a59

Observation 390f0c41-6019-48cc-895d-beb02767722d · outbound

This paper cites Large Language Models for Education: A Survey and Outlook.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Large Language Models for Education: A Survey and Outlook

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:57.992768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:57.992768Z digest=sha256:5319cc95e5f0ce5f051f2d777e99d3b690a175c8535056a720f6aa2a1d73b8d0

Observation aeef9f0d-8016-4b8a-9ebb-db11e3b5e7e7 · outbound

This paper cites Informatics , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Informatics , volume=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:58.067412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:58.067412Z digest=sha256:72b1df348e03b399dea7fed4a1280eea02e44dd0059acc71b358ca71150af5c0

Observation 5ee1f032-12a6-4a78-b728-c3a0956cd609 · outbound

This paper cites Authorea preprints , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Authorea preprints , volume=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:58.133388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:58.133388Z digest=sha256:f42a3eaf32347f4799194b536b8d98e42cff71a19dc0b6da5dfd6f0690d0a0c0

Observation f2d6e1fd-df4c-46c2-b21b-bd20a6992d6d · outbound

This paper cites A Survey on Large Language Models for Code Generation.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions A Survey on Large Language Models for Code Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:58.218018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:58.218018Z digest=sha256:535c78d49c6aff1609a0e3f3c6a46df15af62be2271b4f2e246a101d29a19b71

Observation ce426873-190a-4a13-b211-ac2dc8f6daed · outbound

This paper cites Findings of the association for computational linguistics: ACL 2023 , pages=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Findings of the association for computational linguistics: ACL 2023 , pages=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:58.276090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:58.276090Z digest=sha256:9643a5c65459b6afc9ed64f565e037ba392bab686a077d4728539c6c3bef10bf

Observation daf1f126-3b89-4513-9613-6bd799100796 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:58.373783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:58.373783Z digest=sha256:c590d312982e77b2a7872621bdb7acc6899e06deedbe1615dc9538caa90dd1ea

Observation 9ad09991-2733-4a69-b2d2-1394c071c5d2 · outbound

This paper cites Nature Machine Intelligence , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Nature Machine Intelligence , volume=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:58.432299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:58.432299Z digest=sha256:afd5acc44d63416444b7853f3936a993f42dc8a841af9beb395d40ea8bccb8e2

Observation 748335bd-f319-4e63-88f2-f636cba9aaba · outbound

This paper cites Nature Machine Intelligence , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Nature Machine Intelligence , volume=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:58.533335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:58.533335Z digest=sha256:94009550a4c024acc005ba3b3a591fab213bb402f593852274072eeb5e801cbe

Observation 222983a9-cc05-42b4-9c64-782afc464c71 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:58.636508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:58.636508Z digest=sha256:590c3cdc4f88bc01ebc2f83799eefa917aa798ab859353460bd80d2bc3c1cd56

Observation dbeede4c-e780-46e5-a574-0617323ee4a2 · outbound

This paper cites SafeWork-R1: Coevolving Safety and Intelligence under the AI-45\^.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions SafeWork-R1: Coevolving Safety and Intelligence under the AI-45\^

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:58.717021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:58.717021Z digest=sha256:7770e0b305fa6b612701d6d3e7620cd1b1875f65eb0150f12d9d5bdf0ec7d4f8

Observation 56b24cb1-7064-471a-b02b-00add06844b8 · outbound

This paper cites International AI Safety Report.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions International AI Safety Report

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:58.799632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:58.799632Z digest=sha256:d4fda19994725a565714be7d03100b003d472a11985370aa80ae1dc841f263c7

Observation 5d63aebd-b504-4428-95cc-ce8752d04fa0 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions On the Opportunities and Risks of Foundation Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:58.879692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:58.879692Z digest=sha256:5227cf8440c9037afb604f308321e161e793b96748edc5b9e96e1672de14b730

Observation 1e66d62e-529c-48a2-a58f-e1cdf6c721cb · outbound

This paper cites URL: https://nvlpubs.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions URL: https://nvlpubs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:58.951261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:58.951261Z digest=sha256:0cf043a7761e8736f4a50fb39fcd312c85455b2af6a437fba9d82c0d4991e041

Observation b360ee57-36a2-46cb-a833-06da870baba3 · outbound

This paper cites 2023 , publisher=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2023 , publisher=

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:59.008206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:59.008206Z digest=sha256:125803bd7484245e3d901a86543a6ec614bf0cd49cf22ae25ef57a3836cc0ef0

Observation 5a9c45d4-ddc9-43c1-b98c-8c06f7769717 · outbound

This paper cites Official Journal of the European Union , number =.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Official Journal of the European Union , number =

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:59.080115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:59.080115Z digest=sha256:b8c45bd5ca988971c792073526dd120c11c09a9b4c4bed90eb1f52cc77cf374c

Observation 9be8b5df-975c-4dd8-a686-e3653d6ddfc2 · outbound

This paper cites Concrete Problems in AI Safety.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Concrete Problems in AI Safety

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:59.137059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:59.137059Z digest=sha256:0712c87dab77911005233c6dd08fbad6dff5ef26ab26530ad145c5ed4cf4bd29

Observation 8efbc6e0-74c3-459c-b776-130303cc9bc7 · outbound

This paper cites The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:59.205593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:59.205593Z digest=sha256:d203c422863103fefa2c8a138c3d20bd6b6c79e0ef1f792fbaa0bc5c5dd71eb9

Observation 8a135a9f-d109-49df-a0f2-92f3f064bc7f · outbound

This paper cites 2023 , institution =.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2023 , institution =

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:59.285088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:59.285088Z digest=sha256:83b0dff0e252813a3fa64fe19d3068ef0abc3144c1ec3990106ca4910c4b3e25

Observation 6daa9d79-8d71-45f1-9aa5-e15c3f669de0 · outbound

This paper cites 2024 , institution =.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2024 , institution =

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:59.361896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:59.361896Z digest=sha256:eb9779e25886eb3e3274cf90adf6392a5c482f9f7d38888a52c1f2dc7ebd4431

Observation e9a03aee-bca8-498b-8f50-d160eccb2051 · outbound

This paper cites Advances in neural information processing systems , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Advances in neural information processing systems , volume=

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:59.435844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:59.435844Z digest=sha256:cf5da488eb7bff7b7d75fb5587830da5748c6b5a07004e299c2c198037ed7acd

Observation c47077ed-cc0a-4a6e-ab67-b7481b5059b4 · outbound

This paper cites Learning and individual differences , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Learning and individual differences , volume=

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:59.503198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:59.503198Z digest=sha256:9b600b3c9774a6b332e1c7ff4b280754f7fc4fd43b932f981805e1eedf891e7b

Observation 52cf64d8-cd5a-4dd9-9a07-e7b62e2ba3eb · outbound

This paper cites Nature , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Nature , volume=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:59.565923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:59.565923Z digest=sha256:b42ce9848699fbac60d7379cfc136a190d37ecc6ba9d1003fb539f08688f08c4

Observation ac885b44-caf0-485b-9938-a7ba948372cc · outbound

This paper cites Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP) , pages=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP) , pages=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:59.639381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:59.639381Z digest=sha256:599ebeaab4fd2d0cc48b75288e1076a4dc67ca9b79f070aa0d2894a9168694c5

Observation 47d6f59f-5d31-4646-b257-d3685c254a71 · outbound

This paper cites International conference on machine learning , pages=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions International conference on machine learning , pages=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:59.694477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:59.694477Z digest=sha256:6454b70e2762d7c6695659ea64628543640e11470293adc02f960f248568a43e

Observation 70d3e225-4228-4e19-918b-1865550ba9d4 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Transactions of the Association for Computational Linguistics , volume=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:59.767550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:59.767550Z digest=sha256:249faa921ff68077aa15725762b43396d2a49d1e44c0a418c6727995ef0711dd

Observation 58efd87a-497f-4f7a-898b-ddab44108a5a · outbound

This paper cites Language Models (Mostly) Know What They Know.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Language Models (Mostly) Know What They Know

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:59.849263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:59.849263Z digest=sha256:f86d02afb7cf525008d6d58f51ed042bf843f7dd76f72f1f1547c8651117fad0

Observation d73020be-2075-43d4-ad95-feb6caa8cbdc · outbound

This paper cites Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:59.925943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:59.925943Z digest=sha256:4f21f2c825a6daf4a9a77e6f4d9304251ae8535de16cb5491f4ad684af65aeb3

Observation 458a7d69-d789-4a61-a49b-fe9a906afcfc · outbound

This paper cites British journal of applied science & technology , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions British journal of applied science & technology , volume=

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:59.998324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:59.998324Z digest=sha256:0c28026ba6565ef6686ddd23d4a0d6829ec76f22bea42fb18519a9f00cdde2d9

Observation c43fb9e4-383d-4975-95c1-0bff492d3ee3 · outbound

This paper cites arXiv preprint arXiv:2407.14937 , year=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions arXiv preprint arXiv:2407.14937 , year=

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:00.088094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:00.088094Z digest=sha256:e6c11ce1846b698747771f4fed6f096aa09c627c90ae75d2665bba04c858def0

Observation bfbada33-5f0b-4969-87a0-b359b19dae5a · outbound

This paper cites arXiv preprint arXiv:2507.04446 , year=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions arXiv preprint arXiv:2507.04446 , year=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:00.167017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:00.167017Z digest=sha256:85bcd638ab602b2b46b334fcf8fdfeb0755737f2bb553638456b1532b661cd26

Observation 0ad08a4e-0b2d-4ac9-919b-4fc516602ed0 · outbound

This paper cites GPT-4 Technical Report.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions GPT-4 Technical Report

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:00.240230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:00.240230Z digest=sha256:20ea4ecb90f336688501e9bb8678bdb88fb5da494562d12cbc9a1b5684c25482

Observation 05f7f820-9b95-4e06-9a9c-b267f1078fe3 · outbound

This paper cites 2026 , eprint=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2026 , eprint=

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:00.313906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:00.313906Z digest=sha256:c849f881a9a3b706b4461925ca79292101e414655974de7d42aa64f4ea24da60

Observation cf244df0-50a8-4695-bd2d-cc95f8f3808b · outbound

This paper cites , author=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions , author=

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:00.401032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:00.401032Z digest=sha256:5ea62ffadef26207ae3d3085cd6880686a786efd66f05611403688d9dbb8ca6c

Observation bcdf61de-112d-4b06-b144-6634ddf5c162 · outbound

This paper cites , author=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions , author=

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:00.513630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:00.513630Z digest=sha256:bc753b7dfa4368e5acb9a6c3aaa5a71c4b7277a9ae35c26bfc54c0311de65463

Observation 1e0546cb-6f39-425c-8ec2-c63d1bd82928 · outbound

This paper cites IEEE Transactions on Human-Machine Systems , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions IEEE Transactions on Human-Machine Systems , volume=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:00.587989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:00.587989Z digest=sha256:790b26bde59b270652641d261dc90c036d7491ae4b9f89236e9ade8df51e2060

Observation a7830ed4-96ab-45c9-a058-429e7f379557 · outbound

This paper cites Policy sciences , volume=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Policy sciences , volume=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:00.647582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:00.647582Z digest=sha256:b7f107789e71949b08671f2bc0b868e4d91997c6aee0aceda67e469a48a16f3f

Observation b86df9b2-6a42-4a2d-a482-a313c3bc73ed · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:00.756145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:00.756145Z digest=sha256:eed17e5646526985ec377fd7422222398bc64c02859f1b701ff3e335d95e6057

Observation bffa723a-232d-49f6-8515-ba6ec89e3f6b · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:00.841653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:00.841653Z digest=sha256:9aaaa21a45eb7afb188439de1e835d0a1e4661dbd11a31befc6cd12af890234a

Observation fb3ccb79-5d81-4b5c-b4aa-5eeb3d941e82 · outbound

This paper cites 2025 , eprint=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2025 , eprint=

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:00.925825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:00.925825Z digest=sha256:4da48e3ce962ea563a8e24a1f9674a53fcda6ed2748da911cca10efaf546ef47

Observation c312933b-1446-44d6-9d18-e3fe0bef0a34 · outbound

This paper cites arXiv preprint arXiv:2601.10543 , year=.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions arXiv preprint arXiv:2601.10543 , year=

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:01.007366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:01.007366Z digest=sha256:60fc492b301b1bbfe4f46e03f3681bd497acb11e3781434db0615813f62fa2df

Observation dde8f715-429a-4a7b-a486-671e594e1b93 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:01.084617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:01.084617Z digest=sha256:aa04864cb482ac5c77830993f0aece0d0c36b788d9bcc2f8aed0515682f11323

Pith citing papers

No inbound Pith citation observations are available.