Pith. sign in

Paper Citation Record · LEDGER

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation

As of 9 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 3 inbound Pith citation observations for arXiv:2505.18556.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18556 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:32:31.644019Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T13:18:40.623601Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T17:51:41.885862Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b431b791-0fd2-4ac2-90f8-61f550819cd4 · outbound

This paper cites URL: " 'urlintro :=.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:28.368220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:28.368220Z digest=sha256:1dbcc6ff95732428c3a5e3205239a1669850258ecfa85bdd95bdf9e3e01a3d1e

Observation 3172e170-78cf-451b-a001-86059b6107a0 · outbound

This paper cites write newline.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:28.460572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:28.460572Z digest=sha256:ce4039cdfcce44eee8fb5a4d2faf6ffa0bf0a039ed46ac37d12d74df4d7025b8

Observation e41aa9c5-5712-43fd-88d4-784ad7b7964d · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.729169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:28.546314Z digest=sha256:64d385bcb41222f532807bba63d6b3d5951ea8ccada038a11343d39a766cba58

Observation b126d2e1-42ec-4643-b165-ef19f7956778 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.556435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:28.583911Z digest=sha256:6a64fca341b9b877da8afa60a184cacff5fe18e9c52e84b7a52095677e9a6e79

Observation 09c0e544-7a5d-481a-80b9-12e2c9a810fa · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.407729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:28.655389Z digest=sha256:9fa2ecfceb495f4292b0f3a4914fbc0579f52090fe2a50bbc7cf332c266a353b

Observation 4ee2aaab-b997-4149-aaac-e5d97dde2032 · outbound

This paper cites Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:28.748572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:28.748572Z digest=sha256:a525c861ab57b499389fc7fc1a4ac45fd8caff0c018d62806970f01a5c8ab3f2

Observation 506c0bed-abf0-423f-8d97-e31c7977d43e · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:28.798711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:28.798711Z digest=sha256:7087afb727bc108248dce5ca2f0188b777d4850f525b88d9083fe6c95d768794

Observation 1b0ffcb1-bb92-4f76-a247-4ef24ff1250b · outbound

This paper cites Z-BERT-A: a zero-shot Pipeline for Unknown Intent detection.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Z-BERT-A: a zero-shot Pipeline for Unknown Intent detection

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:32:32.295390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:28.849543Z digest=sha256:3ffaa3bac83d2ec75a35ac4d990cdce54f21bca5fcfdbdb7375e7bd01d9b050d

Observation de1eedfa-e02d-4ed1-9394-bf001ae8372d · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.319299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:28.949530Z digest=sha256:4e5d7ee3c0310f24405d09ed6a20285cb26f6d4a76d73ea668d4839525e406cf

Observation 0c92bd8b-b5ce-4772-aa71-25501a945267 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.258130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:29.030810Z digest=sha256:ce6b355f13bb3ca2dcbcf4f2c94f9adb1da077639bcf6f7dd1159167625a86dc

Observation a2b84897-36d2-4d2b-84ef-9eaac98ee83e · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.074711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.074711Z digest=sha256:24bf4e4753933e4e2d1fff5440d6416a19452b800695183be0519ff647415ac1

Observation 83ece4b4-ec12-41c9-9c8b-c471793193ec · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.133774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.133774Z digest=sha256:c7af79708b4a7934b80240701ef849023819ca1b53d06118ca8dd11bc22ef222

Observation 31a650b5-5135-4391-bedc-1ad88c054921 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.139141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:29.224704Z digest=sha256:e0a2cfae8064a704d1ffff87c4ee497e1581a04331400670f8c2eeb1894c3652

Observation 1a8a5e0b-ea4e-4b82-bd59-b5d088003154 · outbound

This paper cites GPT-4o System Card.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation GPT-4o System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.287134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.287134Z digest=sha256:f674a76c0e98fb9d46923600ad846d8ecb19f33ed4eb256c292276a0bc00ebb0

Observation 882510a1-c0c5-42ec-88dd-476aca8704e5 · outbound

This paper cites OpenAI o1 System Card.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation OpenAI o1 System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.320086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.320086Z digest=sha256:1841fae741643813586365028bfa2f0fd05c836782ca63a4202053caea12c714

Observation 5f64d5f7-277b-49ed-9b32-797732415312 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.432174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.432174Z digest=sha256:f7f1181881405c2a63f4a6f9070b0e7b31239cf4c53c2a2f768596908303e4f7

Observation df9b1b16-5935-4d78-86de-5a319a58074c · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.510906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.510906Z digest=sha256:09ad535af4b42ae1e1bb731bc9f2a0a63f430ff1ae2c89595db39c3cf7bd495f

Observation 51e6f370-574e-401e-8292-f543d35becb0 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.039835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:29.539612Z digest=sha256:13e5baa9977277eaf77775f5257ca2a60b9e993b414c169c409bf8cd31f176ce

Observation 81e35d6b-b051-47dc-b881-75a029c5c93b · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.913348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:29.633976Z digest=sha256:3c4d769775c860b79088de1cc3e980e66e5936303a78b4c820d40d90e38cfff4

Observation 9820e81e-1a97-4abd-9d41-c424982aea51 · outbound

This paper cites Open Sesame! Universal Black Box Jailbreaking of Large Language Models.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.722321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.722321Z digest=sha256:588320f5dcaeebcb924778461b1380e5979ab6101a0034871d20292a8215455d

Observation 3251d485-49a5-4b94-835f-62a66192a6df · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.776349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:29.790056Z digest=sha256:6dd533bc3139d1877441cf4eb716fb7f839695519a24d34053e75ff5d6b37d98

Observation 242c95c8-92d6-4104-b7ba-e7860536e439 · outbound

This paper cites DeepSeek-V3 Technical Report.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation DeepSeek-V3 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.834692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.834692Z digest=sha256:2581ff9897eb7b6bea0bb20f45862c2d4450c9c666693c1dcef0f2ec90db5929

Observation 87f4dd05-014a-4240-89c4-2321628b6b58 · outbound

This paper cites FlipAttack: Jailbreak LLMs via Flipping.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation FlipAttack: Jailbreak LLMs via Flipping

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.915875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.915875Z digest=sha256:50a610fafac28a1bfbdddfb1a3776e2ecf417b139065e26f6ee437fd24fd2419

Observation 0d61b577-7030-4b17-a34d-833315d29f99 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.706179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:29.970460Z digest=sha256:010255234ebdca2f75f8ae7c0aa7aad061845c4bf727e4357ce3a0388fd84cb0

Observation 5991bb0f-c38b-479b-86b3-021d59dccfc8 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.021994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.021994Z digest=sha256:84eb2bbc988a91f0458eeacbb537fff86555ffcef616b8ffd89ae78396a40f61

Observation a8e1c95f-3cdb-4dd7-939e-16de1224f06f · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.064116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.064116Z digest=sha256:001fe81305b6d275007b4dd598416b794904dbed3d17c5339abd50b1f282464c

Observation 4601b9e5-8120-418b-a1fd-0461eb6ece5c · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.577088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.138925Z digest=sha256:57601ace09d9332ba054018ce928f76e0af1693a5a5e8029d8d35ef202107c6d

Observation bd04e60b-495f-4d0b-859a-a2c0b99cf229 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.464284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.229495Z digest=sha256:b8f7c0995313bb69fa39de55e9068ba25bd99f2ac647f93bb112bf9289079dea

Observation 48f16cf2-8790-4d24-b2d3-f2a00cc4603e · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.372707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.267198Z digest=sha256:d42d42ed755af388253b3408c9fcfea59196c95837da92c554db950576fbf472

Observation 7f88d5ee-0bcf-4695-9c52-275fba0d9acb · outbound

This paper cites IntentGPT: Few-shot Intent Discovery with Large Language Models.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation IntentGPT: Few-shot Intent Discovery with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.352898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.352898Z digest=sha256:f6321382ad13ffb415a12f70988bb77237db7c7f214543c676e023e0bbbd750d

Observation 38ce1dea-3956-4aaf-bb0c-193073a598e0 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.218954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.432860Z digest=sha256:edf8b1ae50df910d9f797b460d52dc13bb4160f9e7f855a894a8cb747c13490b

Observation 3d3b302f-04de-4aeb-ad6b-60458330f346 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.141225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.483183Z digest=sha256:da930c2f2d5966ed428a6cc1488a792b1bd841f81aaf915c5c4d8cfaa21bd4fc

Observation 003315bc-28a4-43f6-b6d9-af64111e6ab6 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.031105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.534650Z digest=sha256:456c6c3745f709800d11ef4140d36dcd92513e5c32beb7003dd342f848dbf36d

Observation ef5bb73a-5f1c-4140-ab2f-96bfc7a71cc3 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.904284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.565324Z digest=sha256:60844c78abbcced6b145bf7a4a3d28e57e9c6da1213fef85bdb439b6c07d8c99

Observation 61305ab5-bfa8-431c-a5a8-6c4fadff7ba4 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.820745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.652851Z digest=sha256:7f1c2e2ce80fd45ca25ba38b1a5c610679ead899122e02dfb2081b8d9e058d66

Observation 50dc41cf-1536-4e7b-869a-64ad6d1d0615 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.720380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.720380Z digest=sha256:8c3d62be5e6eef03026b2755ca2ab2bd4edca63fb7e97b6f8e7894d8e825cf59

Observation aaf8c749-40c7-4757-a814-50bc26cf75e1 · outbound

This paper cites RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.773525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.773525Z digest=sha256:974d164f848e848cc233f4fea4239ac052e84d30cb48968708fb68abbb2a77aa

Observation 739e6a25-7f9e-4c5f-9925-4b25d9d16e52 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.698410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.817410Z digest=sha256:ec62a8a16f7ae4b2df3e2af368bec0488b4bc5af9a95fa06297a6c5573a859c0

Observation 213c098c-63ae-4d11-aee0-962ef541656e · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.523536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.859591Z digest=sha256:3512e949c09d5e2798ef1a1e17ade89abb82d298d1350f27a083acec475596f5

Observation 791bdf0f-7e6b-4e1e-bcd8-d243c270fae0 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.919666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.919666Z digest=sha256:176b1ffef6d507a9cc408c34090bc1a4f2e74a19659ea76269cf3fb45047c798

Observation e2fcdcaf-b1b8-419a-b03b-8dc884c67fa1 · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:31.034762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:31.034762Z digest=sha256:bea46196cebdd9ba607d4f5afbcfa8e4e35db7f19374d3ab5ae436a46afc3c78

Observation 041362cf-344d-431d-9f65-b2f644b51495 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.326229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:31.088188Z digest=sha256:f0bb3beee244e4ebc0fbf06bf0fc7190d0b621394ebcda344054745eb798fc68

Observation 4cf9ed77-4cdd-4c78-ba1c-76616ac9bf66 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:31.124752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:31.124752Z digest=sha256:ca3da5b9662924e0fbca62dd4a93915fc383e66ae03f1b6d0e4876e5f33b4a81

Observation deb7a500-333c-4e9e-af3a-52a49d38e0eb · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.130348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:31.185153Z digest=sha256:82cda6256120d07392580b82c2f437224e84c2c5e280ffa106d6766bab01587a

Observation 7549e762-cf0c-45fa-9cd5-6c51e9a911ec · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.012359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:31.272832Z digest=sha256:b6df39cbda3d2b8fdcc7ed86a6acf694830b3314b5a0812797fd207cf6db2964

Observation cafffd9b-ee29-49b4-81c8-a09b699277a3 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:32.862069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:31.346209Z digest=sha256:5b5e73044a666806cab2e651661c6f851c186f86e2147c2f994dc170a91330db

Observation 306dfaa0-926f-4e67-9edb-c8e95cb94cd8 · outbound

This paper cites WordGame: Efficient & Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation WordGame: Efficient & Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:31.393629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:31.393629Z digest=sha256:d1b5d14b18a700d5dc643c043f8c5f646e88d27599627e8e9465154fa7386ee5

Observation 363920a5-7472-43d6-bfb9-c5b14637ff60 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:32.706784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:31.465433Z digest=sha256:1529ee01edfbafed0f6d3cfcac0a111a6f4d28fc311173fa5ae6b13b077e7083

Observation 8c3068a1-aab5-4bc1-8738-4a37c628852f · outbound

This paper cites Autodan: Interpretable gradient-based adversarial attacks on large language models.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Autodan: Interpretable gradient-based adversarial attacks on large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:32.486577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T14:32:31.558726Z digest=sha256:3e7e239bb1a145a5409633303ea7e934f8e710cbf4df9b3ccb61c60d785f8300

Observation 06bbe084-a062-490a-81fd-b9a1ac80b7ae · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:31.644019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:31.644019Z digest=sha256:c415b55381c76e8990c9fcd04a251068aeba5d8fdeb65fd815730b71c9d165da

Pith citing papers

Observation 4a48bd81-4ad5-4d9e-b690-0bbcb0aab7c9 · inbound

Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain cites this paper.

Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:51:41.887953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T17:49:42.112564Z digest=sha256:080e3d0503c8d0115b3cbb780e8b8de4a914e213584991455ee0d3fbd80a5224

Observation 49805767-a29c-4379-acc2-cebec45087db · inbound

Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting cites this paper.

Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation

Reference 10

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T18:08:13.066346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T18:04:44.543311Z digest=sha256:418ed65fa5de5f0e2cb5df36496516c50910ac24710458ef69ad575cbbb0d8de

Observation d40ec3d7-7faa-417d-b17c-e93b52d8736f · inbound

Incomplete Prompt Jailbreaks in Large Language Models cites this paper.

Incomplete Prompt Jailbreaks in Large Language Models Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T13:18:40.623601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:18:40.623601Z digest=sha256:87fb61eadb0853ded033b3eacdcbb652770e15cb0a88ba41b645edf7f81db494