Pith. sign in

Paper Citation Record · LEDGER

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation

As of 10 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 3 inbound Pith citation observations for arXiv:2505.18556.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18556 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:32:31.644019Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T13:18:40.623601Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T17:51:41.885862Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b431b791-0fd2-4ac2-90f8-61f550819cd4 · outbound

This paper cites URL: " 'urlintro :=.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:28.368220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:28.368220Z digest=sha256:1dbcc6ff95732428c3a5e3205239a1669850258ecfa85bdd95bdf9e3e01a3d1e

Observation 3172e170-78cf-451b-a001-86059b6107a0 · outbound

This paper cites write newline.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:28.460572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:28.460572Z digest=sha256:ce4039cdfcce44eee8fb5a4d2faf6ffa0bf0a039ed46ac37d12d74df4d7025b8

Observation e41aa9c5-5712-43fd-88d4-784ad7b7964d · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.729169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:28.546314Z digest=sha256:2b92ed9ed697e61c1346d65f81b8d8b3110f6600b70e5d3fd661bbbde0e5f419

Observation b126d2e1-42ec-4643-b165-ef19f7956778 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.556435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:28.583911Z digest=sha256:afe801cd01c276278ced79eeb0b19e3b31241882487c876fb9b842cb5b0172a9

Observation 09c0e544-7a5d-481a-80b9-12e2c9a810fa · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.407729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:28.655389Z digest=sha256:b898e366c55dce82baf9fd228095a22771adee07317306d1d89ad7288bc320d8

Observation 4ee2aaab-b997-4149-aaac-e5d97dde2032 · outbound

This paper cites Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:28.748572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:28.748572Z digest=sha256:a525c861ab57b499389fc7fc1a4ac45fd8caff0c018d62806970f01a5c8ab3f2

Observation 506c0bed-abf0-423f-8d97-e31c7977d43e · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:28.798711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:28.798711Z digest=sha256:7087afb727bc108248dce5ca2f0188b777d4850f525b88d9083fe6c95d768794

Observation 1b0ffcb1-bb92-4f76-a247-4ef24ff1250b · outbound

This paper cites Z-BERT-A: a zero-shot Pipeline for Unknown Intent detection.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Z-BERT-A: a zero-shot Pipeline for Unknown Intent detection

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:32:32.295390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:28.849543Z digest=sha256:dc3312d9d4723c6e54680cb042fbbaf8b9ba04c9eaee4702d289880b3278033d

Observation de1eedfa-e02d-4ed1-9394-bf001ae8372d · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.319299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:28.949530Z digest=sha256:c7844e3949dafe5139fd1601a9f4c7dd4d8fc6abcba423a193fd8d38f0aca00c

Observation 0c92bd8b-b5ce-4772-aa71-25501a945267 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.258130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:29.030810Z digest=sha256:0d509146eb192bbe8db9ac2116aa11bd98d63744573ba3702343edb1f7af5708

Observation a2b84897-36d2-4d2b-84ef-9eaac98ee83e · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.074711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.074711Z digest=sha256:24bf4e4753933e4e2d1fff5440d6416a19452b800695183be0519ff647415ac1

Observation 83ece4b4-ec12-41c9-9c8b-c471793193ec · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.133774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.133774Z digest=sha256:c7af79708b4a7934b80240701ef849023819ca1b53d06118ca8dd11bc22ef222

Observation 31a650b5-5135-4391-bedc-1ad88c054921 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.139141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:29.224704Z digest=sha256:78683dcf31e6cd4eed8feb5bc3cf24e6abf86d518e1fe8a9317cd30875cecf5f

Observation 1a8a5e0b-ea4e-4b82-bd59-b5d088003154 · outbound

This paper cites GPT-4o System Card.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation GPT-4o System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.287134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.287134Z digest=sha256:f674a76c0e98fb9d46923600ad846d8ecb19f33ed4eb256c292276a0bc00ebb0

Observation 882510a1-c0c5-42ec-88dd-476aca8704e5 · outbound

This paper cites OpenAI o1 System Card.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation OpenAI o1 System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.320086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.320086Z digest=sha256:1841fae741643813586365028bfa2f0fd05c836782ca63a4202053caea12c714

Observation 5f64d5f7-277b-49ed-9b32-797732415312 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.432174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.432174Z digest=sha256:f7f1181881405c2a63f4a6f9070b0e7b31239cf4c53c2a2f768596908303e4f7

Observation df9b1b16-5935-4d78-86de-5a319a58074c · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.510906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.510906Z digest=sha256:09ad535af4b42ae1e1bb731bc9f2a0a63f430ff1ae2c89595db39c3cf7bd495f

Observation 51e6f370-574e-401e-8292-f543d35becb0 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.039835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:29.539612Z digest=sha256:2211496406282de3b3d4f618b7d915ecb592eb469a7acc058439228432cb9279

Observation 81e35d6b-b051-47dc-b881-75a029c5c93b · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.913348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:29.633976Z digest=sha256:054a2c5b4d141fcdda3648cd0fcf95759096de4b9ff19b4a5e27924b07c728d0

Observation 9820e81e-1a97-4abd-9d41-c424982aea51 · outbound

This paper cites Open Sesame! Universal Black Box Jailbreaking of Large Language Models.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.722321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.722321Z digest=sha256:588320f5dcaeebcb924778461b1380e5979ab6101a0034871d20292a8215455d

Observation 3251d485-49a5-4b94-835f-62a66192a6df · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.776349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:29.790056Z digest=sha256:4c53748d1ae2ddd3cfb3cb0c18066776469b80c6cfe38ce5d47a5636982d7fe6

Observation 242c95c8-92d6-4104-b7ba-e7860536e439 · outbound

This paper cites DeepSeek-V3 Technical Report.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation DeepSeek-V3 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.834692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.834692Z digest=sha256:b7a057f80cb9cc829cea29c7b0b7804365b5b40af4c477a0c857e7536b950fcd

Observation 87f4dd05-014a-4240-89c4-2321628b6b58 · outbound

This paper cites FlipAttack: Jailbreak LLMs via Flipping.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation FlipAttack: Jailbreak LLMs via Flipping

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.915875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.915875Z digest=sha256:50a610fafac28a1bfbdddfb1a3776e2ecf417b139065e26f6ee437fd24fd2419

Observation 0d61b577-7030-4b17-a34d-833315d29f99 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.706179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:29.970460Z digest=sha256:f4b31d98e342ee45014085a901a61e1a744dae2b3dbee2f58b203d56cb8b5bb5

Observation 5991bb0f-c38b-479b-86b3-021d59dccfc8 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.021994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.021994Z digest=sha256:84eb2bbc988a91f0458eeacbb537fff86555ffcef616b8ffd89ae78396a40f61

Observation a8e1c95f-3cdb-4dd7-939e-16de1224f06f · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.064116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.064116Z digest=sha256:001fe81305b6d275007b4dd598416b794904dbed3d17c5339abd50b1f282464c

Observation 4601b9e5-8120-418b-a1fd-0461eb6ece5c · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.577088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.138925Z digest=sha256:9b6b46e388f214cdeacaa3dce3816f9176b1df5cb3f68c97818086cacd94dcae

Observation bd04e60b-495f-4d0b-859a-a2c0b99cf229 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.464284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.229495Z digest=sha256:1e44f402134abfc6c09ad6b23740542aea6c53e8c4b39e001245baeddbb611c8

Observation 48f16cf2-8790-4d24-b2d3-f2a00cc4603e · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.372707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.267198Z digest=sha256:8dae449b6f8e7bf061aaaf353c338ff2d7f20f29fb280642ee2af7db4c6d344a

Observation 7f88d5ee-0bcf-4695-9c52-275fba0d9acb · outbound

This paper cites IntentGPT: Few-shot Intent Discovery with Large Language Models.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation IntentGPT: Few-shot Intent Discovery with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.352898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.352898Z digest=sha256:f6321382ad13ffb415a12f70988bb77237db7c7f214543c676e023e0bbbd750d

Observation 38ce1dea-3956-4aaf-bb0c-193073a598e0 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.218954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.432860Z digest=sha256:7cd00b17e29b2b9760724e0b8025d9cdf01c700bfef944cbed22edfdcaf03061

Observation 3d3b302f-04de-4aeb-ad6b-60458330f346 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.141225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.483183Z digest=sha256:ea4dd7fb69d2c66bb9800fb8f112bcfb8951f3730e4749980e499bd096f3289d

Observation 003315bc-28a4-43f6-b6d9-af64111e6ab6 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.031105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.534650Z digest=sha256:fd993c6367b60ae7a5719396aedf789901cb59c325a82d3dec8b135b19cee0d3

Observation ef5bb73a-5f1c-4140-ab2f-96bfc7a71cc3 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.904284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.565324Z digest=sha256:c9254e333721893f8508a4ede72bb9452f67a6dbc8ef0090511bf5084f4b5e55

Observation 61305ab5-bfa8-431c-a5a8-6c4fadff7ba4 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.820745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.652851Z digest=sha256:282356d271c7d337a498167faf2ee189ee983b57d5f285e1fcc774f50c083dc9

Observation 50dc41cf-1536-4e7b-869a-64ad6d1d0615 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.720380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.720380Z digest=sha256:8c3d62be5e6eef03026b2755ca2ab2bd4edca63fb7e97b6f8e7894d8e825cf59

Observation aaf8c749-40c7-4757-a814-50bc26cf75e1 · outbound

This paper cites RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.773525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.773525Z digest=sha256:6dfd4c947661da6a0d3b347558c53c6f4b0eab571e460a245da52b1604a4b315

Observation 739e6a25-7f9e-4c5f-9925-4b25d9d16e52 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.698410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.817410Z digest=sha256:199c56f6a6ef55e47a092af99c95b0c33af9320640f45d00ced5906b5e894e5d

Observation 213c098c-63ae-4d11-aee0-962ef541656e · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.523536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.859591Z digest=sha256:2ec662738cdef8c1b0f9e8409caa983bf6511a42f98d1174f429463ad9e98ffc

Observation 791bdf0f-7e6b-4e1e-bcd8-d243c270fae0 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.919666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.919666Z digest=sha256:176b1ffef6d507a9cc408c34090bc1a4f2e74a19659ea76269cf3fb45047c798

Observation e2fcdcaf-b1b8-419a-b03b-8dc884c67fa1 · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:31.034762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:31.034762Z digest=sha256:bea46196cebdd9ba607d4f5afbcfa8e4e35db7f19374d3ab5ae436a46afc3c78

Observation 041362cf-344d-431d-9f65-b2f644b51495 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.326229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:31.088188Z digest=sha256:4624381067940bd4c4ed1309da2a5c7794f0f0e0264b5ffc3f21e1b0b6de3f73

Observation 4cf9ed77-4cdd-4c78-ba1c-76616ac9bf66 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:31.124752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:31.124752Z digest=sha256:ca3da5b9662924e0fbca62dd4a93915fc383e66ae03f1b6d0e4876e5f33b4a81

Observation deb7a500-333c-4e9e-af3a-52a49d38e0eb · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.130348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:31.185153Z digest=sha256:fb84ee49e6b90f54d2d9444f27647e4838e09ea5f2266f1cc3ca3acb4681300c

Observation 7549e762-cf0c-45fa-9cd5-6c51e9a911ec · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.012359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:31.272832Z digest=sha256:725172007a2ce736a0a29b25f3402bda3eb42ae606781a1891ea4a12a7f5d6e1

Observation cafffd9b-ee29-49b4-81c8-a09b699277a3 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:32.862069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:31.346209Z digest=sha256:82d5fe1def57e9a61b8d9b3b37fa2fac235b8838bc7f1e3715e102ce48f36f0f

Observation 306dfaa0-926f-4e67-9edb-c8e95cb94cd8 · outbound

This paper cites WordGame: Efficient & Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation WordGame: Efficient & Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:31.393629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:31.393629Z digest=sha256:d1b5d14b18a700d5dc643c043f8c5f646e88d27599627e8e9465154fa7386ee5

Observation 363920a5-7472-43d6-bfb9-c5b14637ff60 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:32.706784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:31.465433Z digest=sha256:7b380ec0caf0129b427b0a79dc982f715a46c28eead10470a4c6b271d98281bd

Observation 8c3068a1-aab5-4bc1-8738-4a37c628852f · outbound

This paper cites Autodan: Interpretable gradient-based adversarial attacks on large language models.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Autodan: Interpretable gradient-based adversarial attacks on large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:32.486577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:32:31.558726Z digest=sha256:342afa8cb0e83e81b5e416e41380f5bf1b34c6154fb6c4755e142d274c410864

Observation 06bbe084-a062-490a-81fd-b9a1ac80b7ae · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:31.644019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:31.644019Z digest=sha256:c415b55381c76e8990c9fcd04a251068aeba5d8fdeb65fd815730b71c9d165da

Pith citing papers

Observation 4a48bd81-4ad5-4d9e-b690-0bbcb0aab7c9 · inbound

Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain cites this paper.

Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:51:41.887953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T17:49:42.112564Z digest=sha256:82492401a8fbecd570aefa72394cf49b2581c8288a8e9405ec418ec938e883c1

Observation 49805767-a29c-4379-acc2-cebec45087db · inbound

Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting cites this paper.

Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation

Reference 10

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T18:08:13.066346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T18:04:44.543311Z digest=sha256:889cc0507a8ab582806d6a9e30272e58a8ec50d5ae6d6a6ff1790ed1cb6b7862

Observation d40ec3d7-7faa-417d-b17c-e93b52d8736f · inbound

Incomplete Prompt Jailbreaks in Large Language Models cites this paper.

Incomplete Prompt Jailbreaks in Large Language Models Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T13:18:40.623601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:18:40.623601Z digest=sha256:87fb61eadb0853ded033b3eacdcbb652770e15cb0a88ba41b645edf7f81db494