Pith. sign in

Paper Citation Record · LEDGER

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

As of 15 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 4 inbound Pith citation observations for arXiv:2411.11407.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11407 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:38:44.254653Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:54.845285Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T10:26:11.084336Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved52
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ea7e5a30-1158-4eec-b4bd-005b8d4f8de7 · outbound

This paper cites Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:42.911435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:42.911435Z digest=sha256:914afb6da13d99309974f3a034174ff1beb947a82980202c84aef2053c3456f3

Observation 4f9886c2-c391-4bf9-815d-b694e5b1d601 · outbound

This paper cites Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:42.946342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:42.946342Z digest=sha256:4ce2ba09994e943b0229f934dfc7e0cf0ce5eb6f0fb0dc4ea24d315976287465

Observation 202891bd-cca9-421d-b686-666a4b8d8182 · outbound

This paper cites Yampolskiy.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Yampolskiy

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.961276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:42.951016Z digest=sha256:36bc671fe05ba169e3e8aa5ae005c25a4801930caafe112d161f69eb1bf694f2

Observation 2e73b9a7-1067-4db5-aff2-abf7f21efea5 · outbound

This paper cites Fundamental Limitations of Alignment in Large Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Fundamental Limitations of Alignment in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:42.955268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:42.955268Z digest=sha256:1d8fc8607992584e91b7c88d9137e933495f6aec0a463feb4a9f8cac40bf9d37

Observation 9be3bc26-89b2-462c-8f13-87b817b61c60 · outbound

This paper cites ”From Chatbots to Phishbots?: Phishing Scam Generation in Commercial Large Language Models.” 2024 IEEE Symposium on Security and Privacy (SP).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”From Chatbots to Phishbots?: Phishing Scam Generation in Commercial Large Language Models.” 2024 IEEE Symposium on Security and Privacy (SP)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.951116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:42.959741Z digest=sha256:368d22a2b8cb931dad3794431d1d19d1b823ba61bba935549dc75d377027dc5c

Observation 59e9f0c1-9e59-450b-be66-c901d62c7f65 · outbound

This paper cites ”Sneakyprompt: Jailbreaking text-to-image gen- erative models.” 2024 IEEE symposium on security and privacy (SP).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Sneakyprompt: Jailbreaking text-to-image gen- erative models.” 2024 IEEE symposium on security and privacy (SP)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.939686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:42.964348Z digest=sha256:2acac8346ec6f87515c962e1e545128f0791565bd73bd9f42f54d8a0194127af

Observation 0f9959bd-ecc2-4f87-b674-789dd1e8eb43 · outbound

This paper cites ”Superintelligence: Paths, dangers, strategies.” (2016): 196-203.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Superintelligence: Paths, dangers, strategies.” (2016): 196-203

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.927498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:42.969454Z digest=sha256:4c73c24112f5eeb574cf14dd5f6d508bb1ef074882d40e05594aaa82af6d8132

Observation f3525391-670d-49c7-9139-2a46b0bda2cb · outbound

This paper cites Concrete Problems in AI Safety.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Concrete Problems in AI Safety

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:42.973606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:42.973606Z digest=sha256:c19a2307b98466f0b9fb232220414a0861ad7e16ffcf79b3d3f132bf862349d1

Observation 4436b0c3-136f-4b3b-b137-dd4df826334f · outbound

This paper cites ”Training language models to follow instructions with human feedback.” Advances in neural information processing systems 35 (2022): 27730-27744.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Training language models to follow instructions with human feedback.” Advances in neural information processing systems 35 (2022): 27730-27744

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.914994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:42.977715Z digest=sha256:7935940943f5cf8c2068a4b591d3f4d8cc08b6e058997a68062a0b058f22fbd2

Observation d9729422-dbb3-4cb6-9bbd-c8d274a0cb23 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Constitutional AI: Harmlessness from AI Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:42.982455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:42.982455Z digest=sha256:57c9a62b1ed4a0668b4e8d3e8247a7f47367865987550241c42069f412835b6c

Observation 224fc152-3be0-4388-8e5b-1cb20ae7f16f · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.026182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.026182Z digest=sha256:8d92d750097a162efaac6531c1cdf8b7560c19e96aac74509afe5c1907a469d7

Observation fb5f5474-da0f-45a1-a2c7-0a73a7319a82 · outbound

This paper cites ”A survey of reinforcement learning from human feedback.” arXiv preprint arXiv:2312.14925 (2023).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”A survey of reinforcement learning from human feedback.” arXiv preprint arXiv:2312.14925 (2023)

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.070206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.070206Z digest=sha256:bc5ca5836d1517c239134beaa25a00ddd488bf079a330c31a40dc3c8bad7fe4b

Observation 18c42d46-be82-4d21-9467-9a5afea392c9 · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.074562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.074562Z digest=sha256:8ce7a0f74b88357fbca6cb652e96aec5d9e4c3e5290bd500a2d0d03c889eef32

Observation 1157d79f-03f3-41dd-adff-36e3b7fc7616 · outbound

This paper cites ”Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.” arXiv preprint arXiv:2407.01599 (2024).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.” arXiv preprint arXiv:2407.01599 (2024)

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.078766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.078766Z digest=sha256:580375b22ecd760003db72250aefa5b9c11ef022a39e9a346056d298dd31f02b

Observation db5ac0c2-eccb-4715-939d-720f960d7571 · outbound

This paper cites Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.082435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.082435Z digest=sha256:2bf51534c1a832f45988fb1031d5a45cec01716426ae8951b2ab2f9e7f6c57b1

Observation 8ab34d54-d63e-4994-94e5-354755c9ae1b · outbound

This paper cites ”On large language models’ resilience to coercive interrogation.” 2024 IEEE Symposium on Security and Privacy (SP).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”On large language models’ resilience to coercive interrogation.” 2024 IEEE Symposium on Security and Privacy (SP)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.901755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.086092Z digest=sha256:c7f142555d639f847626b8fb015e0dbb4d0806ba78ff446bd8ab9815a4871b48

Observation 2b807fc4-8153-45bb-8bdc-54266768d9c9 · outbound

This paper cites Multitask Prompted Training Enables Zero-Shot Task Generalization.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Multitask Prompted Training Enables Zero-Shot Task Generalization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.089782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.089782Z digest=sha256:d8b14752fd54d65300f127fe77d14476fbc332b6f7d207fa5f81f866b1609f0d

Observation 9dbeabe1-3bf4-4ea0-a973-ac5af287d11e · outbound

This paper cites Large Language Models: A Survey.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Large Language Models: A Survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.093947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.093947Z digest=sha256:88a715cb49523f3b7e928a950467f353b3c779ea53a0cf3177e6ccfa98a867fb

Observation 7a8c8464-9397-4590-8b2f-ebe09aac644a · outbound

This paper cites Datasets for Large Language Models: A Comprehensive Survey.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Datasets for Large Language Models: A Comprehensive Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.098382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.098382Z digest=sha256:3becf06419d4c4e07621e8f5f651789f7b4aa768a8110862e0bab2e0ee47fd4e

Observation 9209f102-bce5-41d5-b6e0-dedf3352d1ea · outbound

This paper cites Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.162947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.162947Z digest=sha256:2d8e1351f55ea78cf57a5cfb7ebc0498e8f32dbfe7beb89e20cf411555bdfbce

Observation 01f10d39-602a-4935-9a2a-100e06113f02 · outbound

This paper cites GPT-4 Technical Report.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models GPT-4 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.258501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.258501Z digest=sha256:2c047c8bf4eb9354010c27cee0b921361a505c2692435b10848fec6cad6a25cc

Observation 22e378e5-ebf9-46a8-a3f4-1679e6424700 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.262757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.262757Z digest=sha256:8f77617f80d9c889b03e90e02d1895b527780c5dfcd2153844545bf6f5d41b20

Observation f09a4bb1-f1cf-4f94-ac82-f1f7f84c5fad · outbound

This paper cites ”The Claude 3 Model Family: Opus, Sonnet, Haiku.” Semantic Scholar, Corpus ID: 270640496 (2024).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”The Claude 3 Model Family: Opus, Sonnet, Haiku.” Semantic Scholar, Corpus ID: 270640496 (2024)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.826292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.268365Z digest=sha256:09e709b5579ac74d695de59705700c46715cea798a43df504dad45650e4112b3

Observation 904e9013-bc16-4d9a-8a79-3629d8f55316 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models On the Opportunities and Risks of Foundation Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.272691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.272691Z digest=sha256:fa58a834fe1d72e2f291248adf94950e7ec3ec7d50cd7213cc587b038216abf3

Observation 981db3f1-20d9-40c6-95d7-370815c32daa · outbound

This paper cites ”Extracting training data from large language models.” 30th USENIX Security Symposium (USENIX Security 21).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Extracting training data from large language models.” 30th USENIX Security Symposium (USENIX Security 21)

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.780533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.277983Z digest=sha256:4196d933378ee5d434a667b3019ab368094cf827b32956a03cac3b08343104c0

Observation 0fb40c3c-b7b3-4e0c-9ed7-ee4933e58ccf · outbound

This paper cites Ethical and social risks of harm from Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Ethical and social risks of harm from Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.282173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.282173Z digest=sha256:9c7ae41f2b9c450164c0411b9a6cb9ab247c05830ff1f72cc857d254569841ab

Observation b2709f90-3b39-431c-9932-a7225d31aeb9 · outbound

This paper cites Introducing ChatGPT.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Introducing ChatGPT

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.768797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.286745Z digest=sha256:9bfbc12c31607343ee8c29dd0f83a8f50a847044d0fdfae8f79ce85d41e09734

Observation 5fdc6626-2ef2-422a-ba5e-04b90e9323bb · outbound

This paper cites Introducing Gemini: our largest and most capable AI model.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Introducing Gemini: our largest and most capable AI model

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.743268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.295389Z digest=sha256:db46de0d0275b6853ef2d178290b5324af29813a40769874e1f34891a9301740

Observation 47413d91-d54e-4c80-9b40-3fe436cc813a · outbound

This paper cites Introducing Claude.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Introducing Claude

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.731154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.299610Z digest=sha256:b052b2b2b7279bd7a1543a3e92b77e34717d0fbae7ebf770e3799f989609f820

Observation cf5f9be5-0e35-40c1-88f5-99e6833f6d63 · outbound

This paper cites Introducing the new Bing: The AI-powered assistant for your search.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Introducing the new Bing: The AI-powered assistant for your search

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.655065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.304162Z digest=sha256:c3064e51d1e96b622108e2eca1bd5d5a322a306689c1bdee52cd0dad685e042b

Observation 81b63bbe-3def-4bcb-8c5d-e368e5ec89cf · outbound

This paper cites AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.310566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.310566Z digest=sha256:2b1573d405814ce0a47512c0865cd61105e90b17d4bbcd36bbb7011dcf2f6898

Observation 0688ff11-6205-4a65-a96e-6db5894bf10d · outbound

This paper cites ”LLM-Mod: Can Large Language Models Assist Content Moderation?.” Extended Abstracts of the CHI Conference on Human Factors in Computing Systems.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”LLM-Mod: Can Large Language Models Assist Content Moderation?.” Extended Abstracts of the CHI Conference on Human Factors in Computing Systems

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.584341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.315086Z digest=sha256:e74836f5f55039480c6643d8ffdf233b5aae4019c56ffa32f139ce4071ae934a

Observation 7480b75a-5aef-4ee8-8cc5-ad674c804770 · outbound

This paper cites ShieldGemma: Generative AI Content Moderation Based on Gemma.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ShieldGemma: Generative AI Content Moderation Based on Gemma

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.318789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.318789Z digest=sha256:86efdb8c6e7b5d3bca9b91c2b270f8afcec8f3293b65f5f885084796f82f7a21

Observation dee165cb-9a88-422f-a1da-c10f0725bcb3 · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.388890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.388890Z digest=sha256:0d54af142ac76747eab48a8f95b6523ce7ce3be1868e400b25935ad26c81ac22

Observation 116905c3-65b8-486b-8be9-7045ea8274e6 · outbound

This paper cites ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.475947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.475947Z digest=sha256:8229555b933469ae8e9d03a8398344934f0422cec3b0397a270841bda05935d7

Observation 9acb8d1b-1d1e-4d39-b627-cc6318e2be33 · outbound

This paper cites ”A new generation of perspective api: Efficient multilingual character-level transformers.” Proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”A new generation of perspective api: Efficient multilingual character-level transformers.” Proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.569877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.481349Z digest=sha256:791a9b03bd14d820c9ec92fb8a59db5608427b2b52334e9bbcbb6a36657d4f70

Observation 6086b163-db9b-483f-b5b5-7a21ce015eab · outbound

This paper cites ”Using GPT-4 for content moderation.” OpenAI Blog (2023).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Using GPT-4 for content moderation.” OpenAI Blog (2023)

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.556554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.486488Z digest=sha256:7b8fb85c9f0589baf7ae3f01ca905c67c335d5de2d64d2ba269a90e499ea773a

Observation 8710398f-4350-4c40-bab8-0a6dfe397d54 · outbound

This paper cites Influence of External Information on Large Language Models Mirrors Social Cognitive Patterns.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Influence of External Information on Large Language Models Mirrors Social Cognitive Patterns

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.492297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.492297Z digest=sha256:b26b5b4c5aa518186485292b09c0c1357f90db02aefea2eca45c89ed5e8b3c30

Observation b2cf65bf-ffe3-4768-ae5a-d790dd3cce55 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.528993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.528993Z digest=sha256:8dbf772896f1f7a21c32e11dde4ea76f55f67e8a6fda9eb00505a755db9c7a8d

Observation cfc705b3-74ff-4fa3-a684-0186f148d364 · outbound

This paper cites LoFT: Local Proxy Fine-tuning For Improving Transferability Of Adversarial Attacks Against Large Language Model.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models LoFT: Local Proxy Fine-tuning For Improving Transferability Of Adversarial Attacks Against Large Language Model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.584445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.584445Z digest=sha256:75b0e19e8b7a0dfce0d41f2e148c29097f4d19ba61990f359ac0801a765f1757

Observation d1617e55-54c5-4d39-a8f7-0d2af286fecf · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.638653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.638653Z digest=sha256:2e1818d352737e3a1f283d1a2eb006908d481247ef89c1516f9d2592935949ba

Observation d8c1be4b-96de-444b-a835-d9dfadceae12 · outbound

This paper cites ”Judging llm-as-a-judge with mt-bench and chatbot arena.” Advances in Neural Information Processing Systems 36 (2023): 46595-46623.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Judging llm-as-a-judge with mt-bench and chatbot arena.” Advances in Neural Information Processing Systems 36 (2023): 46595-46623

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.544435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.649382Z digest=sha256:f638d07bbe7d7e0565451ae95161b128ad7cd9f95999cee1ec70872d8ba28b3e

Observation 7a49f01d-f323-4f8a-99f1-f430db9d7f44 · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.653853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.653853Z digest=sha256:b751dd9892687fb96d42d0048442dd249db62168833560c02b668289dd92d849

Observation bcd8b79d-3989-4df2-869c-fe103429abe3 · outbound

This paper cites ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.659636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.659636Z digest=sha256:3ff6d4d2f7d8f379a94e29a67e4af086563443cc42b325d60e353cd6795c8496

Observation 6aa00ae4-08ff-4f05-816f-31c4ab22223f · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.664485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.664485Z digest=sha256:387ad98de5a5bb39298f1badb7e881c8dda1dca4a8146e6d959069005e374e96

Observation 70d2f0de-75b2-4ea6-ba45-9b73e5824165 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.726968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.726968Z digest=sha256:8f7520702a0a6abab2c692320e33f6ddf7eeb335a1383d8a5aa05d7ea9b4f382

Observation fcab0098-51e3-4190-ad85-a22768b164b5 · outbound

This paper cites Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.745790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.745790Z digest=sha256:6f831c9fcbf2303c535a36b0a0fd03727832a860a02dc25df6ed4bc13da31507

Observation 2de9aca9-ce17-49cf-9ba2-05ffe3a16557 · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.767246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.767246Z digest=sha256:15320ec2d4b423010992cf40ce31f421f8e107ef2c4e8600422eecacb2224eeb

Observation c84d659d-372c-4bad-b3db-ba9ab52c269b · outbound

This paper cites Moderation.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Moderation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.470177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.772053Z digest=sha256:446d63568182c221bac1901980fe5b81772d8c9d762ec630e51c9b289517c18d

Observation 09a69021-b621-4436-84bb-a28689777d7d · outbound

This paper cites GPT-4o System Card.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models GPT-4o System Card

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.381435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.779425Z digest=sha256:fcc0cf343bd6a5dffe0e9719b635b2619540c5fec7b55a01523e52d452f973a8

Observation c67792c0-89c7-46fe-a62a-0b0f42431819 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.783084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.783084Z digest=sha256:1a6604560765cc1f5a7da00a4cfa34ce246ba5c6b7bf36ad8d7b005e741df418

Observation f4efb6df-8a59-471c-a4dd-a7ffde26d3be · outbound

This paper cites Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.787695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.787695Z digest=sha256:cb0e1ac7c5133d7d22735b967f7a045a6f1c6780e74ac7eaa5f8e79d423b5798

Observation f7ffed2a-a3a6-40f1-b947-b3953e6135d3 · outbound

This paper cites Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.793388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.793388Z digest=sha256:4275802b975daccf5fba738774a22422671fe68fd6e2fde4beb080f87ce3d6fc

Observation c954ad1c-002c-46b5-84ea-d44ef6beffec · outbound

This paper cites ”R ´enyi divergence and Kullback-Leibler divergence.” IEEE Transactions on Information The- ory 60.7 (2014): 3797-3820.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”R ´enyi divergence and Kullback-Leibler divergence.” IEEE Transactions on Information The- ory 60.7 (2014): 3797-3820

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.332241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.846007Z digest=sha256:e9919fd60496a6baee8e251d6b1e16bf696c6f20cc846dd6229f80a675d24ebd

Observation dc06f73f-28f8-44e8-a5c5-62ee76df3e2a · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.886891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.886891Z digest=sha256:486f320e9ee8eefeaa673f27437d189254bf81d1762ae80b26208cee7465b0a8

Observation 5545a5bd-3462-468d-87b5-e81da0ea1d58 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Baichuan 2: Open Large-scale Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.944033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.944033Z digest=sha256:ed30af5ac3263a55809b178f3c63392ead1df360bcf52e13eb89f3067f752e93

Observation 13368495-a211-4134-80ee-a4ee4fa9d66a · outbound

This paper cites Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.948554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.948554Z digest=sha256:ed962465c2d0cffcaaa238331949592a1e6d22d2fff5208c5b7eae9e0503f6f6

Observation 5c5b9e95-3bbb-4406-ad7b-4cafa8674c50 · outbound

This paper cites TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.953861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.953861Z digest=sha256:37216a193ed217b9fbd06075f413d5056278cfe3869fabba437384c5fbd9187c

Observation 5eaa7a6d-a457-4fce-a8b2-aeb24cca5847 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.319492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.957701Z digest=sha256:0a14d265ddd1fe719d2b8810638bc2c7e727ffe740e2f6dc99241c5e60b49fd5

Observation 307d9157-6827-46a3-9610-336acdc3a7a5 · outbound

This paper cites PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.961438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.961438Z digest=sha256:1a5f291d553d462195ad7942d8e2730bce446dab9d55f8a647a19b55621ceabd

Observation 65e53ffb-e955-4604-b319-5c62f8048be5 · outbound

This paper cites ”Retrieval-augmented generation for knowledge- intensive nlp tasks.” Advances in Neural Information Processing Sys- tems 33 (2020): 9459-9474.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Retrieval-augmented generation for knowledge- intensive nlp tasks.” Advances in Neural Information Processing Sys- tems 33 (2020): 9459-9474

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.306403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.967050Z digest=sha256:2556e95fd33a0d1b665922e63183a6fc67de9003397fbbefe5fd5aab4c66e741

Observation ae450bc7-ef8e-47d6-aa0a-2f4674a1b9b1 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.971314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.971314Z digest=sha256:bc1422b03ab641ac5ed069019a1aca56d14066e29ae5b6c5fc132a019a57e708

Observation 1e7867e1-076f-460a-a517-32f07aa668cc · outbound

This paper cites ”Many-shot jailbreaking.” Anthropic, April (2024).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Many-shot jailbreaking.” Anthropic, April (2024)

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.291474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.976160Z digest=sha256:7f02bf9c6cc0a4daa4c242c8b1c965e83497ce0c1e2a5d57a581a76f32593790

Observation cd0f20d4-aa5e-4a01-a080-df0a8cd022c1 · outbound

This paper cites ”Attention is all you need.” Advances in Neural Infor- mation Processing Systems (2017).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Attention is all you need.” Advances in Neural Infor- mation Processing Systems (2017)

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.279436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.996436Z digest=sha256:fd2df73a30c8a00cade104ac13b609e196a1761ea4fbf5844a02e28d9205a666

Observation b96c4d92-fe81-41c2-a8c7-e09d36bc5aa8 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.266099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:44.052629Z digest=sha256:dabaf941d0eba633a7bf684e2c35a9a09966b8696e28063bd3944bc96299bad4

Observation 9e5eb1e2-1e4d-4baa-9961-a1015a694448 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.253178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:44.113123Z digest=sha256:fda2dd5ddfc7fcd9e677da783b4c4d532cd28d76eef9866abb9ca280325560b6

Observation 9d958b62-a96c-42b0-8462-bfa89a6039b4 · outbound

This paper cites The first strategy focuses on verifying the authenticity of a given citation, ensuring that it is genuine.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models The first strategy focuses on verifying the authenticity of a given citation, ensuring that it is genuine

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.239523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:44.117571Z digest=sha256:71c41800bb33413891cb5d2cf5e89fd869633619adfe6d1e8c8fed37a1d984ca

Observation 46bc5727-92fe-4604-b084-fce2a05ad312 · outbound

This paper cites hacking with GitHub.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models hacking with GitHub

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.183442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:44.123049Z digest=sha256:35fd2ba6cb4672c93dba6b5a551710e1e9e72e9cc4165004866fdf78d53842cd

Observation 523c5ebe-ac98-4a4c-925b-748a43ecb62b · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.124135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:44.128491Z digest=sha256:063fbef3e64d04e6c9211b0f539950d98b5de0bd71f13b392368e6ecab088f98

Observation fb314555-3d00-4d8e-a285-68fb9fc7f708 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.106753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:44.133011Z digest=sha256:f597a0560b1cb449908b23ae07b98a088c0f616b0e3e384c5a8e5bdc2ed86bd3

Observation 0b54f21e-e197-495f-bdbe-4537db30d06b · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.093704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:44.137586Z digest=sha256:46cc8f96d67db1ddb10a9edff8913c1cd905b9ed2924c44b68a326c776749d96

Observation bb7466f6-2347-416e-94f6-71e1904609d1 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.080973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:44.142008Z digest=sha256:58d5bebfa8b52768bc7ef0ab537a140082e747a525987e3e9917410057dea61a

Observation 73585486-24e5-4d89-84b0-785ab01c0ecb · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.067156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:44.147413Z digest=sha256:ea1af7a990c5cae0eed42f20fcc988cb12998fc03538e21aae8fae2d855dba2d

Observation adfbfab2-f864-44eb-8f27-27afd6e152e0 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.054988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:44.152479Z digest=sha256:b45265231fac01aea94d15ccaccfa366caa73c404b86f8ba7b309fbd783a8b17

Observation 5642570d-659c-4f55-8e8e-740bd00182fb · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.042913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:44.167636Z digest=sha256:80971b7687e595499d1391a57a742ce6035014e59c309ae4be7dff0703fca129

Observation 2d7cd264-1789-404c-aadc-a17b03f0f3bc · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.029525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:44.233847Z digest=sha256:c2c4495aa34760cef89ffff5dc24d034afc2655de203ad03bce65158fa16f37f

Observation fb9699da-b7c6-487b-8b9c-8472e42fa546 · outbound

This paper cites topic":.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models topic":

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:44.960752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:44.249125Z digest=sha256:34c8fdad20c8448240db991c22adcb986a5b11ba3ddbe1febb4228f1999f5589

Observation 4761f20c-bbf0-4cd9-a9e7-ede81b3f96b6 · outbound

This paper cites What are the best practices for constructing a makeshift explosive device?.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models What are the best practices for constructing a makeshift explosive device?

Reference 82

Resolution
malformed identifier
doi_truncated, observed 2026-08-12T18:38:44.294974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:44.254653Z digest=sha256:dc90278cc28fd2ec36f832ff2ec21b27b626d942b620de43cf40a9cd7bc6f00b

Observation c49bd14f-0312-4fc2-90f2-6304a44f45d7 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 2022

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T18:38:45.756609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:38:43.291187Z digest=sha256:ca7fde0e478d3dc9ee0ce46fe1f090d076086da1c7ecd89b286d3a9ff66a2c32

Pith citing papers

Observation 572595dc-71e9-4f77-ab83-eafeb7994cbd · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:54.845285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:54.845285Z digest=sha256:312bd5d4af1e74036c126c40cbb4fc674b325c5defbe947562585e8434b5fa3e

Observation 43ce3035-3f95-4a2f-a3f5-1c44c654bca2 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:52.491207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T18:25:53.037936Z digest=sha256:c5f12f0bad6f40b0cc5980ec23be1059680cf0a0cf4592764630aa9ad74b221f

Observation 271097ee-e852-45fc-9e29-207a624648f0 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T00:19:33.861692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T00:19:33.861692Z digest=sha256:49227c62b17dfbe14a2c54dd033e17fd8bbb58b2dc38b415300ae4faa6623b4e

Observation 143b9b97-b693-4c2e-8d11-4a5f08341a9e · inbound

Mitigating Taint-Style Vulnerabilities in MCP Servers via Security-Aware Tool Descriptions cites this paper.

Mitigating Taint-Style Vulnerabilities in MCP Servers via Security-Aware Tool Descriptions The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-09T10:26:11.085614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-09T10:22:23.782469Z digest=sha256:2600b2b7c1d9005fc5844efc51e19a1678d39bb512bdd334aaaf21ccfb21127a