Pith. sign in

Paper Citation Record · LEDGER

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations

As of 10 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2507.06043.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06043 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:18:47.341663Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:08:04.673125Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T21:08:06.388576Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 20ab7275-4286-4557-9065-dbc99fd0f870 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.488602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.488602Z digest=sha256:fa9653e28fb4038918f0b9608a7f6ed1b8234fe0bd4d1c740ef533e9d7f225cf

Observation dca1a89b-0b75-4bdf-9f69-555eb9834939 · outbound

This paper cites Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.524938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.524938Z digest=sha256:b44240f456643a907a5435d664ecdc5d7b07f683aafe9c01da36e4555fc14de8

Observation 05791a07-f465-41f7-bb98-f0b59ef38b68 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.645786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.645786Z digest=sha256:a9abaefe7a4d43d2e06a214e47331a88b23770adc2a5c11b5adb69a913497a0e

Observation de9f5339-3746-4348-96d9-378f28bdd6ab · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.749367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.749367Z digest=sha256:fb0afdeb5a75e5561f5d3aa323b42085890dd6d8fec204eccf4d5ec98c46fe22

Observation f211f1dd-e80a-4b48-9029-1116493f6a2a · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.849506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.849506Z digest=sha256:7c588f225f961f65e233376dc3435eee112f8f83b33555f10827b8ee9026b9d4

Observation 155f482f-eb0c-4b09-930b-c5236f1f9e56 · outbound

This paper cites JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.941390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.941390Z digest=sha256:45137d8210dd6e30668e4c419a65d9677a114990ddd7ee942ac4bcf553c16f51

Observation c9a0a10f-545e-408e-a876-ba2746afa4df · outbound

This paper cites Recent Advances in Attack and Defense Approaches of Large Language Models.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Recent Advances in Attack and Defense Approaches of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:43.023171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:43.023171Z digest=sha256:596b7d7ce5f1af52cf6dcf3cd07f7f8776cb6536845d7372e6b9df88b86a5094

Observation c2213e70-c872-4395-afc6-7cd63728a498 · outbound

This paper cites DeepSeek-V3 Technical Report.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations DeepSeek-V3 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:43.106007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:43.106007Z digest=sha256:232f60252fac3ac020b82875c08927811c02385187a6ea8f25938279011c6bce

Observation 66a459e4-4b2b-4c9f-9f7c-cc657cb21a57 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:43.155123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:43.155123Z digest=sha256:5e2cae61ec22f5e1070749a9b76eb2847a18716a06909443bb374f1c723fdf9f

Observation 90cbfe8c-bc9c-412d-934d-4bf66a9f80dd · outbound

This paper cites The Llama 3 Herd of Models.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:43.324924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:43.324924Z digest=sha256:eec5f8089394c8367e42f092c79ce27dfffa87ef2d2f63551b685d6769cefbd6

Observation 2a34fef4-bc96-4448-b0cc-ee6154146cec · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:49.761528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:18:43.494982Z digest=sha256:678e76b177e1277a671046a49aa516f913ec1abaa52bae1cd8b3a6a1dbfa1c37

Observation 14825689-1475-4b1b-92ab-944630b02ac0 · outbound

This paper cites Cai, James Wexler, Fernanda B.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Cai, James Wexler, Fernanda B

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:18:49.602687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:18:43.629083Z digest=sha256:22913144e3da539eb9a480521410056367efa2ae7f96cf84415370b046a7abe9

Observation 2baaf5bd-1d05-4f98-852d-db0b98c6f7f7 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:49.424805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:18:43.749206Z digest=sha256:0a1f2ebe6ee3ac627ed36e01daaff3b0228c6554499a76a80be5bfcb98cbd8cc

Observation ee87e432-a3ea-48b7-8126-f346b8db9b00 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:49.237439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:18:43.839587Z digest=sha256:d2d8c27685108ec1ada6d35fe3f3c6fe1a865ecf31937d39279448cbc35798ff

Observation f4e61275-a567-406e-9d36-772ed10f644d · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:49.039356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:18:43.957717Z digest=sha256:bde6e431cd5c5951df6fb797b7ea1e82b0250e5068afac72a01b9d749493b581

Observation ffdae9da-ff4b-4453-b3b6-53c2e2961b41 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:44.110225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:44.110225Z digest=sha256:f87be8305a6fd0f347452fd4042db7f8522424964213735994aaf43ded6693f5

Observation 60fcaa8e-59eb-4e78-9b0e-88e058a1dce9 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 17

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T19:18:48.293208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:18:44.315764Z digest=sha256:c9863a476cd560a014232be5f24b0d9597762349bbcf959a824f92d632b8382a

Observation 92ec32f8-01f8-42ac-9f86-3b9e0d7f35e5 · outbound

This paper cites Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:44.434891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:44.434891Z digest=sha256:9817de9e8ba895a44843b35b027728f43c023782412c60ccbe0d3447fc35c142

Observation e75817f1-eb76-4a4d-ad33-e6183d1af717 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:48.877206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:18:44.557865Z digest=sha256:c826600e55630e50267852bd03eb12c7b0799f9421db19f93e2347f4fca6f864

Observation aa26c873-6fc4-4540-aa2b-924dd1d18413 · outbound

This paper cites GPT-4 Technical Report.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations GPT-4 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:44.729046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:44.729046Z digest=sha256:7cd84b038a1b4feed26800c2363f6d5993fb816b9c22d24a1fae8308558d50d6

Observation a67b6306-0fed-4213-9137-12a769f00d0f · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:44.823274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:44.823274Z digest=sha256:190c081abf942248314bd7ec5977020cd239e500600afe17799479a1a420bfb6

Observation 79a238d2-a2d6-4640-99b8-806ff726846c · outbound

This paper cites Qwen2.5 Technical Report.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Qwen2.5 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:44.931351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:44.931351Z digest=sha256:0dfe969eb880ce3ce1a87539bc9578c87cb12309ef7b804f8a05183a5d98b865

Observation 39ceda5e-54fd-49f1-883c-711e3adea093 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.033818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.033818Z digest=sha256:2c05349ddb4bd0c38449c68bea215993c4bb5593e476248e46a578e6f4a95ec9

Observation f26780fa-9390-45b8-bf79-0ac77c907ffc · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.148774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.148774Z digest=sha256:6b07dd7c87f128dbf491e036bc9e6717d11d71278f3b5ac12f0e9a24f32a9681

Observation b578c107-79d4-4f86-b42f-35891a32855f · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.243985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.243985Z digest=sha256:eac8a627a80c61163a1ee63741d16fe343715e072b4a8a044ec67af76c8da228

Observation b1715f98-6308-4bfa-a4e3-663672797a83 · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.379143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.379143Z digest=sha256:e4463b08f072b966ddaa7d6a1840215351d2dedc12f575698a643d6949959bc8

Observation 5284303c-4ce3-43b0-a30c-41b934161202 · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations A StrongREJECT for Empty Jailbreaks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.500292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.500292Z digest=sha256:b27b16b1bb7fb8e228ab4854edd9d0125d0e2b29a05cc49c4be4cd61d6849105

Observation 8d8f141d-59c6-49ae-ba2b-348faa9e9ffc · outbound

This paper cites Hashimoto.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Hashimoto

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.634496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.634496Z digest=sha256:ea9951952cd8fba559de1bb81bbfed8a3414ce95d45da12c57bffacd523237ba

Observation 758c82ee-0cc8-4b74-a19b-d0aaa8d1712a · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.728426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.728426Z digest=sha256:b91a3448ce4eeee076abad50766a9889f4d942b77b5f0318c4268a4196c8153a

Observation bce1bd3a-1c87-4860-93ae-a6b77a32d2df · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Jailbroken: How Does LLM Safety Training Fail?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.849258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.849258Z digest=sha256:790412c5b7baf2f092a521cfc274e7f2f8ddd624cc20f467899d613bd63298e3

Observation 0a191a1b-a55e-4e05-8ab0-ed4870e40c1e · outbound

This paper cites Dai, and Quoc V Le.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Dai, and Quoc V Le

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.952671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.952671Z digest=sha256:43ac6d9e0a37ab8b495e93b6db75a5a269f6efdce5f735f30c2b2f37eacc37d0

Observation 8f7b9bf5-41b1-4bcc-b490-cae65e440cd3 · outbound

This paper cites Ethical and social risks of harm from Language Models.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Ethical and social risks of harm from Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.047757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.047757Z digest=sha256:ce2fea453b68bfe880cdce463c65ef29443049fd507e5b752c81d013430a11f2

Observation dd2c71f7-cde5-4944-bc27-36a99e3ed40e · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.137678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.137678Z digest=sha256:1e5e015066f021eeef5eac476d0a5c6e7f305b135d78cce40340d01bab499c2e

Observation 65c38dfd-c68d-406b-834d-8f98bd11a601 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:48.613526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:18:46.223411Z digest=sha256:3cd32619537ce4794f5ab94135303c8f5140a6e390c1288b76f0d2c1b57689c9

Observation 2bca0902-16d6-4263-b795-1ee6f7ec13bb · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.295508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.295508Z digest=sha256:14843ff799595a4fab62102fefe9aba53786e8af5ee93e33e1220564e9fece78

Observation 84632e61-765a-4cf8-b5b5-9911eab0c50f · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.428000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.428000Z digest=sha256:54092179638aa9e512eddbca0f4a60946cf1cfc427999a5417ef9f153b5d8a2f

Observation 70640b84-e213-4601-b2c0-aa0b955aa74c · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Low-Resource Languages Jailbreak GPT-4

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.549467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.549467Z digest=sha256:3e5e4017771c7d642daf8546d06e3958cc4df6a9e6213f29c1a69c60fee4deef

Observation 87b57c8e-c12e-49b7-a269-950a04fd0d57 · outbound

This paper cites How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.662354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.662354Z digest=sha256:be1ce253dbaf2b94180b82a44ef547b0e0a8086106fb313f4aa76744564c4b41

Observation 36385948-cc43-4a26-8b82-35896d8f38de · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.786950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.786950Z digest=sha256:61cb4066153acd62de276e0bac8d9b6963096c5b9ccf5d4902257ebf0bfc418e

Observation 66a5804c-728e-4412-a414-3aaff9b71d58 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.864156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.864156Z digest=sha256:b557f19f9ef0e2755677e9a90476aa076e571aa9467082464022b635e357ae53

Observation 0f3429be-0dbf-4ac4-8253-5162e656584a · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Representation Engineering: A Top-Down Approach to AI Transparency

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.970519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.970519Z digest=sha256:e591a1db53556d046905664f11084381d68c4ddd5877a270fa9537c2d11ede25

Observation 822a0e55-2997-4ae1-8622-df97a4992f59 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:47.094379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:47.094379Z digest=sha256:87844aa2bd6e4f1ee47197077a9ba3ded8bc339879b48b5767fc7e922e8e3d8b

Observation 86dd1667-75f1-4985-ba75-b247edf29fbc · outbound

This paper cites online" 'onlinestring :=.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations online" 'onlinestring :=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:47.235990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:47.235990Z digest=sha256:9c176ef2ec0182a4f60ee82d310298a7e3b40c08267939a5c3c0231ce7a71bb2

Observation 13d529b9-ac61-4a1e-8e94-43b2e793f351 · outbound

This paper cites write newline.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations write newline

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:47.341663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:47.341663Z digest=sha256:c693d9256723cf6b91c394833925955619cac7261dec5b0da9985fffcbe82b64

Pith citing papers

Observation 7c1330a7-e8ce-46f9-b491-83cbe88844a7 · inbound

NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs cites this paper.

NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:08:06.426230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T21:08:04.673125Z digest=sha256:4789a0cfb4f5cde6c7d7c5c6942920468efdf1f2c5e78e3b77c3f98ecd78a4da