Pith. sign in

Paper Citation Record · LEDGER

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations

As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2507.06043.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06043 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:18:47.341663Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:08:04.673125Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T21:08:06.388576Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 20ab7275-4286-4557-9065-dbc99fd0f870 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.488602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.488602Z digest=sha256:94a4cd8ad506c2c45a3071f734a3da625e7e8675e1f942e248c9fb974f15c5d2

Observation dca1a89b-0b75-4bdf-9f69-555eb9834939 · outbound

This paper cites Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.524938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.524938Z digest=sha256:55333e6c67b337bb82d38c4bfffaae197c17260240b6ba5058f23063a23e4633

Observation 05791a07-f465-41f7-bb98-f0b59ef38b68 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.645786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.645786Z digest=sha256:8bb5eb77d23c8454d9dedf65eb12e62e5b63c0d1b30d226234deaad04187dc5d

Observation de9f5339-3746-4348-96d9-378f28bdd6ab · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.749367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.749367Z digest=sha256:518dc804d5f160b642c317c9bf773f05a2c597adf2410f1d6ab839b83d7a87ea

Observation f211f1dd-e80a-4b48-9029-1116493f6a2a · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.849506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.849506Z digest=sha256:af02723acf314f8aa90d5b7a80179851bc0a7dcac492e37929b7016dc36a143d

Observation 155f482f-eb0c-4b09-930b-c5236f1f9e56 · outbound

This paper cites JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.941390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.941390Z digest=sha256:bf6d80efab6dc2c7330e8682a60ea40434eea4df315df35dd9c212c97fb53666

Observation c9a0a10f-545e-408e-a876-ba2746afa4df · outbound

This paper cites Recent Advances in Attack and Defense Approaches of Large Language Models.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Recent Advances in Attack and Defense Approaches of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:43.023171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:43.023171Z digest=sha256:e3cd0a9213fa4e0178834f5bfdd672f4d52909505bd38c94c3fa86f45b64cca2

Observation c2213e70-c872-4395-afc6-7cd63728a498 · outbound

This paper cites DeepSeek-V3 Technical Report.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations DeepSeek-V3 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:43.106007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:43.106007Z digest=sha256:c4f974ba98676c95004ae4672dcf7f6989d3c5a53a82244685d8ddc4feb1670b

Observation 66a459e4-4b2b-4c9f-9f7c-cc657cb21a57 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:43.155123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:43.155123Z digest=sha256:6e72754431ab5bbfd3072000b15b36d309b15c06d04d484aaf2c9af01d44b63b

Observation 90cbfe8c-bc9c-412d-934d-4bf66a9f80dd · outbound

This paper cites The Llama 3 Herd of Models.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:43.324924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:43.324924Z digest=sha256:8fad99ed47bb527d01fe515c77a8f3815340f9fb4835e7f98eae1cb9bd16e556

Observation 2a34fef4-bc96-4448-b0cc-ee6154146cec · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:49.761528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:18:43.494982Z digest=sha256:c23ba410b57debb9901ef8d6fbad8eb8ca7ffb9d6418111a76979183fdc42877

Observation 14825689-1475-4b1b-92ab-944630b02ac0 · outbound

This paper cites Cai, James Wexler, Fernanda B.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Cai, James Wexler, Fernanda B

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:18:49.602687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:18:43.629083Z digest=sha256:02ec6dcc18048dc95d21e46041f22ff6702ae07e8759080a5d96fb4a8236e6f1

Observation 2baaf5bd-1d05-4f98-852d-db0b98c6f7f7 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:49.424805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:18:43.749206Z digest=sha256:3d9195597835ba5daf8d6066f9392e1a2792e5c3cd3f8ef0bb3dd0d2683ad72d

Observation ee87e432-a3ea-48b7-8126-f346b8db9b00 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:49.237439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:18:43.839587Z digest=sha256:85d9002add299365436e06ad9e5550e9bc5e73debf1088586ef57b913251cea2

Observation f4e61275-a567-406e-9d36-772ed10f644d · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:49.039356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:18:43.957717Z digest=sha256:1a6fd9c63a911779e8280e046994c2f813bf78d95ae1bfbd2e8218b484e4b7d6

Observation ffdae9da-ff4b-4453-b3b6-53c2e2961b41 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:44.110225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:44.110225Z digest=sha256:4bb1f5d0d6f8f394c80a1a89e7c49161ab1001bea22330c25093112b483b4f39

Observation 60fcaa8e-59eb-4e78-9b0e-88e058a1dce9 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 17

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T19:18:48.293208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:18:44.315764Z digest=sha256:50409007554a0b12a3b8bbbb3ea0308b23ddd8d8ec3c7f98ff84f639461cd238

Observation 92ec32f8-01f8-42ac-9f86-3b9e0d7f35e5 · outbound

This paper cites Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:44.434891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:44.434891Z digest=sha256:51d12e6265339beb96cc2a974045195741e4073620fd7b2ef52e674e29d7fb6d

Observation e75817f1-eb76-4a4d-ad33-e6183d1af717 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:48.877206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:18:44.557865Z digest=sha256:caa3db65892d207d38082af3a637deef2ea9fb1e4474ce93d4dc2d3d8b0994ea

Observation aa26c873-6fc4-4540-aa2b-924dd1d18413 · outbound

This paper cites GPT-4 Technical Report.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations GPT-4 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:44.729046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:44.729046Z digest=sha256:72338ca975d9884929a0c8e118f13a902b3dbf66178e1c6b6151c5d12b3c44d9

Observation a67b6306-0fed-4213-9137-12a769f00d0f · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:44.823274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:44.823274Z digest=sha256:8eb87d2241cc47813d2cafba93125686d398b9bd2f9d6e2bc3e157e6afa3f5f3

Observation 79a238d2-a2d6-4640-99b8-806ff726846c · outbound

This paper cites Qwen2.5 Technical Report.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Qwen2.5 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:44.931351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:44.931351Z digest=sha256:3782e8d5a65eb1bc560b16f10ddf65a4ed32cf3915abdcf2eac9af519d1c6cf1

Observation 39ceda5e-54fd-49f1-883c-711e3adea093 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.033818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.033818Z digest=sha256:268302ee62c8e5930559f26e2f4cc935409414b0147cffc8f17338f917fdf375

Observation f26780fa-9390-45b8-bf79-0ac77c907ffc · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.148774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.148774Z digest=sha256:8d65ca0592ffaf85841897ef83cf153b2401ebcc64a4325adca3aef23342a4b1

Observation b578c107-79d4-4f86-b42f-35891a32855f · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.243985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.243985Z digest=sha256:e377eb165eaadd23553477be7c706d01f88226a32634b56f905e491a0876cbf8

Observation b1715f98-6308-4bfa-a4e3-663672797a83 · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.379143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.379143Z digest=sha256:2866becd233d0b908a27ceef1fc4291ba19d61d74bf013a8ae260e4f83fb5a03

Observation 5284303c-4ce3-43b0-a30c-41b934161202 · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations A StrongREJECT for Empty Jailbreaks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.500292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.500292Z digest=sha256:66d1f66e37d77a41158c9358d288300482bf9e86a0e18347523204483a805aba

Observation 8d8f141d-59c6-49ae-ba2b-348faa9e9ffc · outbound

This paper cites Hashimoto.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Hashimoto

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.634496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.634496Z digest=sha256:de1577904d009452e190ac32404a1d899a8afaa1406ece3e47bda18fbcb7190b

Observation 758c82ee-0cc8-4b74-a19b-d0aaa8d1712a · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.728426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.728426Z digest=sha256:6e0959d997fa16c5b8165a499aa145b94254231bc876a186e13d733b70a1d34b

Observation bce1bd3a-1c87-4860-93ae-a6b77a32d2df · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Jailbroken: How Does LLM Safety Training Fail?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.849258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.849258Z digest=sha256:bd15c8fc33ee706666fa59df94327b2e7701101cfaedffeda0dfaeca3b310558

Observation 0a191a1b-a55e-4e05-8ab0-ed4870e40c1e · outbound

This paper cites Dai, and Quoc V Le.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Dai, and Quoc V Le

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.952671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.952671Z digest=sha256:2343a89907300cbadbae6760aae7fb2a692deab09b99527daeb7efd8b24e10b9

Observation 8f7b9bf5-41b1-4bcc-b490-cae65e440cd3 · outbound

This paper cites Ethical and social risks of harm from Language Models.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Ethical and social risks of harm from Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.047757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.047757Z digest=sha256:9f2637e9fcfa2768001f001765031d23fc1f4a0e0a13ed991bbf95835627b1df

Observation dd2c71f7-cde5-4944-bc27-36a99e3ed40e · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.137678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.137678Z digest=sha256:08171b7eb230c6f375eddc348ee3b105534ba4e6de04ec24750d09502f93335b

Observation 65c38dfd-c68d-406b-834d-8f98bd11a601 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:48.613526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:18:46.223411Z digest=sha256:5b9d8c2aed5db7032618af329bba2a9f4a408502ad270d8fba7cd92875bbfe1d

Observation 2bca0902-16d6-4263-b795-1ee6f7ec13bb · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.295508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.295508Z digest=sha256:050ea71ebb20a66c702056cd1bf3615e19a724f5de76064f1dcd19375a38ddf2

Observation 84632e61-765a-4cf8-b5b5-9911eab0c50f · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.428000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.428000Z digest=sha256:2dd422839f5b41adf059e6c917ed46bee3f82a73b45068b7eca5fedd59d9e082

Observation 70640b84-e213-4601-b2c0-aa0b955aa74c · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Low-Resource Languages Jailbreak GPT-4

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.549467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.549467Z digest=sha256:28171b6f98b517de6bdd17327af56490864199973a63b75daa58eb5c2fccea5a

Observation 87b57c8e-c12e-49b7-a269-950a04fd0d57 · outbound

This paper cites How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.662354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.662354Z digest=sha256:b52abc109b3575c82591b84649188d27db7e90f13f96d71a47fa5af19657fc4a

Observation 36385948-cc43-4a26-8b82-35896d8f38de · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.786950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.786950Z digest=sha256:07a8d5cbf9a87049a71da48473e09fe27712e4f9f9089f3cac90928166178b9d

Observation 66a5804c-728e-4412-a414-3aaff9b71d58 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.864156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.864156Z digest=sha256:7d34819d3a83d051412415e330812e172e1a0e49bd3e46f377840e4c0754f05c

Observation 0f3429be-0dbf-4ac4-8253-5162e656584a · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Representation Engineering: A Top-Down Approach to AI Transparency

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.970519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.970519Z digest=sha256:90f941fe335a10fe7fde5bd3d28b318d483297eb42376a9e00e9c168e04fe05f

Observation 822a0e55-2997-4ae1-8622-df97a4992f59 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:47.094379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:47.094379Z digest=sha256:048f489e0959988224bf2aeea2662d6aef6e4180db1c5503f9198cd3a45a2ce9

Observation 86dd1667-75f1-4985-ba75-b247edf29fbc · outbound

This paper cites online" 'onlinestring :=.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations online" 'onlinestring :=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:47.235990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:47.235990Z digest=sha256:42f2e444cf7cb913af26b52303cdc1dd3253be3b285290c8bd5216fc5a0e74e1

Observation 13d529b9-ac61-4a1e-8e94-43b2e793f351 · outbound

This paper cites write newline.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations write newline

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:47.341663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:47.341663Z digest=sha256:a19a8ef2b230be5049920914da1aa2ffa79e6c8ff4ea438713dec26f06b34dc8

Pith citing papers

Observation 7c1330a7-e8ce-46f9-b491-83cbe88844a7 · inbound

NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs cites this paper.

NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:08:06.426230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T21:08:04.673125Z digest=sha256:b047b16d6be720cb98199b09a9d1fe728489615504ac80674613c82523aedef4