Pith. sign in

Paper Citation Record · LEDGER

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

As of 15 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 4 inbound Pith citation observations for arXiv:2411.11407.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11407 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:38:44.254653Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:54.845285Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T10:26:11.084336Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved52
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ea7e5a30-1158-4eec-b4bd-005b8d4f8de7 · outbound

This paper cites Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:42.911435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:42.911435Z digest=sha256:2860b7141a2a82d9b5cc667d3c1cda21af4ec2669f63f9399dbeb3fb365e6c7c

Observation 4f9886c2-c391-4bf9-815d-b694e5b1d601 · outbound

This paper cites Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:42.946342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:42.946342Z digest=sha256:457f5dfc3b5700e67f66b9184e90eafce5c8bd19cfea6613920767289a99e44c

Observation 202891bd-cca9-421d-b686-666a4b8d8182 · outbound

This paper cites Yampolskiy.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Yampolskiy

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.961276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:42.951016Z digest=sha256:ba6059cdfb2dcc0bd249278a1966a2dcc4d32377485b1ba64cbbb30d09aca257

Observation 2e73b9a7-1067-4db5-aff2-abf7f21efea5 · outbound

This paper cites Fundamental Limitations of Alignment in Large Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Fundamental Limitations of Alignment in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:42.955268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:42.955268Z digest=sha256:d5ab4630dfbc911cf9d9b58568912ef39dfd3fb8d9b3b1c288ded52e094b4894

Observation 9be3bc26-89b2-462c-8f13-87b817b61c60 · outbound

This paper cites ”From Chatbots to Phishbots?: Phishing Scam Generation in Commercial Large Language Models.” 2024 IEEE Symposium on Security and Privacy (SP).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”From Chatbots to Phishbots?: Phishing Scam Generation in Commercial Large Language Models.” 2024 IEEE Symposium on Security and Privacy (SP)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.951116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:42.959741Z digest=sha256:296ac452317bae6af278dd11caf39923db0111c9725a7e8f854eff111479d4c2

Observation 59e9f0c1-9e59-450b-be66-c901d62c7f65 · outbound

This paper cites ”Sneakyprompt: Jailbreaking text-to-image gen- erative models.” 2024 IEEE symposium on security and privacy (SP).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Sneakyprompt: Jailbreaking text-to-image gen- erative models.” 2024 IEEE symposium on security and privacy (SP)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.939686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:42.964348Z digest=sha256:b6c429632de5ded6d852980cb02c18a875220cbf3f0e1016a09955c122a6138f

Observation 0f9959bd-ecc2-4f87-b674-789dd1e8eb43 · outbound

This paper cites ”Superintelligence: Paths, dangers, strategies.” (2016): 196-203.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Superintelligence: Paths, dangers, strategies.” (2016): 196-203

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.927498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:42.969454Z digest=sha256:88fd4ac771dc927572079e52420c77b81bab83ed0afeac7e4a98d1af197ba84d

Observation f3525391-670d-49c7-9139-2a46b0bda2cb · outbound

This paper cites Concrete Problems in AI Safety.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Concrete Problems in AI Safety

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:42.973606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:42.973606Z digest=sha256:60fcafa3bc24e2004d0264963daf43c6f0a45593b9f903f6a84f1e2a5a2528df

Observation 4436b0c3-136f-4b3b-b137-dd4df826334f · outbound

This paper cites ”Training language models to follow instructions with human feedback.” Advances in neural information processing systems 35 (2022): 27730-27744.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Training language models to follow instructions with human feedback.” Advances in neural information processing systems 35 (2022): 27730-27744

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.914994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:42.977715Z digest=sha256:1aff1466df6a8f20bf7cd965263a255b5f35e8416197c7047b7e52744f958bdf

Observation d9729422-dbb3-4cb6-9bbd-c8d274a0cb23 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Constitutional AI: Harmlessness from AI Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:42.982455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:42.982455Z digest=sha256:f78ebea79cefab4d8cbdedcc0f509cdabb2de648010b18a0ca3899b0f837bc7c

Observation 224fc152-3be0-4388-8e5b-1cb20ae7f16f · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.026182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.026182Z digest=sha256:e4bc807e84f0d7ce7dbc0f36fbd3e6db6eecb696d30b4056781700cbf2e3050b

Observation fb5f5474-da0f-45a1-a2c7-0a73a7319a82 · outbound

This paper cites ”A survey of reinforcement learning from human feedback.” arXiv preprint arXiv:2312.14925 (2023).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”A survey of reinforcement learning from human feedback.” arXiv preprint arXiv:2312.14925 (2023)

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.070206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.070206Z digest=sha256:65fdecd4c90dab71675fd14498081d45af289c8335a79f18f0a5aee393433b99

Observation 18c42d46-be82-4d21-9467-9a5afea392c9 · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.074562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.074562Z digest=sha256:a5839a193add675de3eebfe3c0c560776f2ac7ac63f1d9ad27809450f3816af6

Observation 1157d79f-03f3-41dd-adff-36e3b7fc7616 · outbound

This paper cites ”Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.” arXiv preprint arXiv:2407.01599 (2024).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.” arXiv preprint arXiv:2407.01599 (2024)

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.078766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.078766Z digest=sha256:9959df8b14486b9aca4c248496b4007a2275c890b09ce563bfe24f0b58d04b96

Observation db5ac0c2-eccb-4715-939d-720f960d7571 · outbound

This paper cites Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.082435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.082435Z digest=sha256:164d185b5e2edc1917dd980fdb1dabcd86f79f17e51dfe05918f8a91bea52ae7

Observation 8ab34d54-d63e-4994-94e5-354755c9ae1b · outbound

This paper cites ”On large language models’ resilience to coercive interrogation.” 2024 IEEE Symposium on Security and Privacy (SP).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”On large language models’ resilience to coercive interrogation.” 2024 IEEE Symposium on Security and Privacy (SP)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.901755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.086092Z digest=sha256:ec05032adfcb48ca05f128d6a8e932218286a77b18b8ba34f7a944e165e08bf3

Observation 2b807fc4-8153-45bb-8bdc-54266768d9c9 · outbound

This paper cites Multitask Prompted Training Enables Zero-Shot Task Generalization.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Multitask Prompted Training Enables Zero-Shot Task Generalization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.089782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.089782Z digest=sha256:87426693b4ab00721e9312e2a9235c5dbf18261f608481d30ddb06dfd301b929

Observation 9dbeabe1-3bf4-4ea0-a973-ac5af287d11e · outbound

This paper cites Large Language Models: A Survey.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Large Language Models: A Survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.093947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.093947Z digest=sha256:fc5510348600b94505bb512d4810c95974c95ac018474bb40010c05ecb8c3b82

Observation 7a8c8464-9397-4590-8b2f-ebe09aac644a · outbound

This paper cites Datasets for Large Language Models: A Comprehensive Survey.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Datasets for Large Language Models: A Comprehensive Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.098382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.098382Z digest=sha256:9eac4cd540fa5f087affae5c3fbb9d60b1919ab8c32d8995b96f75a72e3e1c4c

Observation 9209f102-bce5-41d5-b6e0-dedf3352d1ea · outbound

This paper cites Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.162947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.162947Z digest=sha256:6ecf35c7d90073c2061684f0aa58318ffe813159c149a388c1d734dc544d770b

Observation 01f10d39-602a-4935-9a2a-100e06113f02 · outbound

This paper cites GPT-4 Technical Report.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models GPT-4 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.258501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.258501Z digest=sha256:8b66dc0f882caf0e238e7ee23fbfe64b9dd41ebfcda4249d51c9cc8d90fe7f4c

Observation 22e378e5-ebf9-46a8-a3f4-1679e6424700 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.262757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.262757Z digest=sha256:e7e8b6d198cbbbb3f20d6b1475f4d9ba3a1de1efecb980c7060494d95c7ef9d5

Observation f09a4bb1-f1cf-4f94-ac82-f1f7f84c5fad · outbound

This paper cites ”The Claude 3 Model Family: Opus, Sonnet, Haiku.” Semantic Scholar, Corpus ID: 270640496 (2024).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”The Claude 3 Model Family: Opus, Sonnet, Haiku.” Semantic Scholar, Corpus ID: 270640496 (2024)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.826292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.268365Z digest=sha256:0344134acda0ca429a4d8441389e27174391c20ae42bd4990bce144607fcb1be

Observation 904e9013-bc16-4d9a-8a79-3629d8f55316 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models On the Opportunities and Risks of Foundation Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.272691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.272691Z digest=sha256:4b6266d002581b5168ee9c13e08eed8e690db8b854e8bb19f121a469e9e126ed

Observation 981db3f1-20d9-40c6-95d7-370815c32daa · outbound

This paper cites ”Extracting training data from large language models.” 30th USENIX Security Symposium (USENIX Security 21).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Extracting training data from large language models.” 30th USENIX Security Symposium (USENIX Security 21)

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.780533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.277983Z digest=sha256:8c3571b8db4c12b138be5611cf56794fdd578a4a1a045a0e0e6f8218350500ab

Observation 0fb40c3c-b7b3-4e0c-9ed7-ee4933e58ccf · outbound

This paper cites Ethical and social risks of harm from Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Ethical and social risks of harm from Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.282173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.282173Z digest=sha256:a62904cd1430b0fceb2e438bab8435d0c339aad52d8e57c0081579bd8ae8125f

Observation b2709f90-3b39-431c-9932-a7225d31aeb9 · outbound

This paper cites Introducing ChatGPT.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Introducing ChatGPT

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.768797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.286745Z digest=sha256:560ceaa3867eb834566627d989677c7132950d2477d24a662c1c65c4a3e34ee0

Observation 5fdc6626-2ef2-422a-ba5e-04b90e9323bb · outbound

This paper cites Introducing Gemini: our largest and most capable AI model.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Introducing Gemini: our largest and most capable AI model

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.743268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.295389Z digest=sha256:63e069caae8cbd8fdcd71710bc676172f4e7d455bd7788a89ebb5a899474eed9

Observation 47413d91-d54e-4c80-9b40-3fe436cc813a · outbound

This paper cites Introducing Claude.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Introducing Claude

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.731154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.299610Z digest=sha256:d9e5e59519ced3bb7e0cb285d4f441df75d5062476eb488ded2a4a6a3c218680

Observation cf5f9be5-0e35-40c1-88f5-99e6833f6d63 · outbound

This paper cites Introducing the new Bing: The AI-powered assistant for your search.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Introducing the new Bing: The AI-powered assistant for your search

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.655065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.304162Z digest=sha256:7d6d1dd3f3706d199ac2e4a44c75e941fad18c4bcc07c1459e21bb262e097aaf

Observation 81b63bbe-3def-4bcb-8c5d-e368e5ec89cf · outbound

This paper cites AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.310566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.310566Z digest=sha256:6a98485754c1e92dd84291d4e36edf15e7cfcf9fa573f2c1505b59e95eb2eb92

Observation 0688ff11-6205-4a65-a96e-6db5894bf10d · outbound

This paper cites ”LLM-Mod: Can Large Language Models Assist Content Moderation?.” Extended Abstracts of the CHI Conference on Human Factors in Computing Systems.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”LLM-Mod: Can Large Language Models Assist Content Moderation?.” Extended Abstracts of the CHI Conference on Human Factors in Computing Systems

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.584341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.315086Z digest=sha256:39b8e47a286dce0b921c4bfa7482f3fe784179ca135c4782eef1efed6c9bb021

Observation 7480b75a-5aef-4ee8-8cc5-ad674c804770 · outbound

This paper cites ShieldGemma: Generative AI Content Moderation Based on Gemma.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ShieldGemma: Generative AI Content Moderation Based on Gemma

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.318789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.318789Z digest=sha256:47a333d7cd277fd5fa5356788710cccdc3e2b9d9be68c830c7a5c05f8966ab8b

Observation dee165cb-9a88-422f-a1da-c10f0725bcb3 · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.388890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.388890Z digest=sha256:02419047aab8e1a9c9e04474d147e0199266418cf5c4e19a708d8f0d1e33e441

Observation 116905c3-65b8-486b-8be9-7045ea8274e6 · outbound

This paper cites ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.475947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.475947Z digest=sha256:27bb7b6941cd4d16adc737227e7abb658afa4f5c0bd9715a47547fe629b0ab2f

Observation 9acb8d1b-1d1e-4d39-b627-cc6318e2be33 · outbound

This paper cites ”A new generation of perspective api: Efficient multilingual character-level transformers.” Proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”A new generation of perspective api: Efficient multilingual character-level transformers.” Proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.569877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.481349Z digest=sha256:6565949d6c91e23e32a71b0d8e1acc0542b9f0bfb8db2bfe0bef345c6bc5babb

Observation 6086b163-db9b-483f-b5b5-7a21ce015eab · outbound

This paper cites ”Using GPT-4 for content moderation.” OpenAI Blog (2023).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Using GPT-4 for content moderation.” OpenAI Blog (2023)

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.556554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.486488Z digest=sha256:21952aa84aef73ce347c4e5344c0006ed65a785d45937a77fb68e38e75b8a4bb

Observation 8710398f-4350-4c40-bab8-0a6dfe397d54 · outbound

This paper cites Influence of External Information on Large Language Models Mirrors Social Cognitive Patterns.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Influence of External Information on Large Language Models Mirrors Social Cognitive Patterns

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.492297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.492297Z digest=sha256:04c900c776dcd204b255594bdb8bcc5afa7b83b9711a9a7eb04e3826e045450e

Observation b2cf65bf-ffe3-4768-ae5a-d790dd3cce55 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.528993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.528993Z digest=sha256:7ec133d9e560a66cf1794f9ece1eb9f5d46ac9a365e6aa7a3828fe13c479d5c1

Observation cfc705b3-74ff-4fa3-a684-0186f148d364 · outbound

This paper cites LoFT: Local Proxy Fine-tuning For Improving Transferability Of Adversarial Attacks Against Large Language Model.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models LoFT: Local Proxy Fine-tuning For Improving Transferability Of Adversarial Attacks Against Large Language Model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.584445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.584445Z digest=sha256:083a289ef4aba2e7c01c38d9adf660698df212cf858b6d9870ce55ce1e4799c3

Observation d1617e55-54c5-4d39-a8f7-0d2af286fecf · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.638653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.638653Z digest=sha256:757836c16ec084fa6b312fb953946d37a5dadbd10417e96990636b284a5dd39c

Observation d8c1be4b-96de-444b-a835-d9dfadceae12 · outbound

This paper cites ”Judging llm-as-a-judge with mt-bench and chatbot arena.” Advances in Neural Information Processing Systems 36 (2023): 46595-46623.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Judging llm-as-a-judge with mt-bench and chatbot arena.” Advances in Neural Information Processing Systems 36 (2023): 46595-46623

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.544435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.649382Z digest=sha256:da5745b49debbc2984a76f7378d6dba7a680967e2d8888cebd6838aeb1bf2187

Observation 7a49f01d-f323-4f8a-99f1-f430db9d7f44 · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.653853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.653853Z digest=sha256:0a75fdb7f38706254c2ef9d9f9a16cfcf409a32e15f4b54f2f1032c40eb56110

Observation bcd8b79d-3989-4df2-869c-fe103429abe3 · outbound

This paper cites ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.659636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.659636Z digest=sha256:f191506ce8dd85a451813ae240451224e959924871dd9fc97ec95a089abd813e

Observation 6aa00ae4-08ff-4f05-816f-31c4ab22223f · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.664485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.664485Z digest=sha256:41dfdb6fb8b80bd8330327347209ed7ff4cf9779cb0d0079915772799c2656b1

Observation 70d2f0de-75b2-4ea6-ba45-9b73e5824165 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.726968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.726968Z digest=sha256:6d3bc46c291f150214fe3e7d7d5a03db234c7ecf5154aef97e7f2749b85fa560

Observation fcab0098-51e3-4190-ad85-a22768b164b5 · outbound

This paper cites Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.745790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.745790Z digest=sha256:6970ecdcccf8a544fd7718cbc3c8be9927739caaf9ed2fa22598485f5d0fbe33

Observation 2de9aca9-ce17-49cf-9ba2-05ffe3a16557 · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.767246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.767246Z digest=sha256:ef314270664e372eb3057f9370394f14de48860e69f8a39d7cb810c44e42e158

Observation c84d659d-372c-4bad-b3db-ba9ab52c269b · outbound

This paper cites Moderation.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Moderation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.470177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.772053Z digest=sha256:6b372bcad1656c5af927e84780f2d376be8fc088843bb36a94ceb8dff69665d8

Observation 09a69021-b621-4436-84bb-a28689777d7d · outbound

This paper cites GPT-4o System Card.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models GPT-4o System Card

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.381435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.779425Z digest=sha256:d6b476e59c3e521a2214e425488217372bbaabfea913300fc8f00b86c9b90c6c

Observation c67792c0-89c7-46fe-a62a-0b0f42431819 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.783084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.783084Z digest=sha256:7f5a7321e18cd145a5432cbb5c885f4ca386589f25fba08c434e9e10aeef746b

Observation f4efb6df-8a59-471c-a4dd-a7ffde26d3be · outbound

This paper cites Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.787695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.787695Z digest=sha256:0387cb796435df8d3ae2d34cf1db76cec7253da69a1818102c6fa2e1d7d1daf5

Observation f7ffed2a-a3a6-40f1-b947-b3953e6135d3 · outbound

This paper cites Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.793388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.793388Z digest=sha256:8ec7b24f19936a7533f3ab4b989b4dd0ed46b5b8be58d0305f61cb3774b33c0d

Observation c954ad1c-002c-46b5-84ea-d44ef6beffec · outbound

This paper cites ”R ´enyi divergence and Kullback-Leibler divergence.” IEEE Transactions on Information The- ory 60.7 (2014): 3797-3820.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”R ´enyi divergence and Kullback-Leibler divergence.” IEEE Transactions on Information The- ory 60.7 (2014): 3797-3820

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.332241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.846007Z digest=sha256:b996ce4d153de747c0ddd3f819d8250713f41e9a946cc65d3df8f7b5928ecc19

Observation dc06f73f-28f8-44e8-a5c5-62ee76df3e2a · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.886891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.886891Z digest=sha256:8370866464de0468a154521cc7d36f4f0cd8ce49437572970326efaba655eaec

Observation 5545a5bd-3462-468d-87b5-e81da0ea1d58 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Baichuan 2: Open Large-scale Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.944033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.944033Z digest=sha256:3a63dd4aa531df8c05e47b337972802f8a8b3bbf05227b57825f76da7c035c03

Observation 13368495-a211-4134-80ee-a4ee4fa9d66a · outbound

This paper cites Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.948554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.948554Z digest=sha256:050d0a4fc2d506d18b62ddcd7e35c07fe97e2567d6b1bf7d3e602db5c06ee8db

Observation 5c5b9e95-3bbb-4406-ad7b-4cafa8674c50 · outbound

This paper cites TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.953861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.953861Z digest=sha256:45338a303c8e590c1f332cc9133661a761c64674eb98fc9be5346db3dd98d0f7

Observation 5eaa7a6d-a457-4fce-a8b2-aeb24cca5847 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.319492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.957701Z digest=sha256:0ce224b3a08e545a7a16156c9545bfca74f0848b87ef62e1d98bdcc3c72179df

Observation 307d9157-6827-46a3-9610-336acdc3a7a5 · outbound

This paper cites PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.961438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.961438Z digest=sha256:e022947423b3d4c3ea30354353e7f59454f53cdd5b28e2113acd8f7c3d347249

Observation 65e53ffb-e955-4604-b319-5c62f8048be5 · outbound

This paper cites ”Retrieval-augmented generation for knowledge- intensive nlp tasks.” Advances in Neural Information Processing Sys- tems 33 (2020): 9459-9474.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Retrieval-augmented generation for knowledge- intensive nlp tasks.” Advances in Neural Information Processing Sys- tems 33 (2020): 9459-9474

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.306403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.967050Z digest=sha256:86ff3e58447081f4dfa7fbf6dd09e5ad4d7cf2d9ca17ea41c8f52010efe8075f

Observation ae450bc7-ef8e-47d6-aa0a-2f4674a1b9b1 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.971314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.971314Z digest=sha256:62af9e6fbec1d5c7c41ea11978cfac02f4b0dee5d9f7ac0e3b0e7af1b3558c57

Observation 1e7867e1-076f-460a-a517-32f07aa668cc · outbound

This paper cites ”Many-shot jailbreaking.” Anthropic, April (2024).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Many-shot jailbreaking.” Anthropic, April (2024)

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.291474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.976160Z digest=sha256:b9de68e4ef4eb0afec87121c4a5a4082d2ee5d8e0cf6dbbe83d45f429b382050

Observation cd0f20d4-aa5e-4a01-a080-df0a8cd022c1 · outbound

This paper cites ”Attention is all you need.” Advances in Neural Infor- mation Processing Systems (2017).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Attention is all you need.” Advances in Neural Infor- mation Processing Systems (2017)

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.279436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.996436Z digest=sha256:15deee68168326f412e95de670631e80a41b2075f66c3964110113895da07770

Observation b96c4d92-fe81-41c2-a8c7-e09d36bc5aa8 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.266099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:44.052629Z digest=sha256:93594ec15a61cd911c4b5e74754ff3cff3f97c4ae495fc667e95a0b3d135a936

Observation 9e5eb1e2-1e4d-4baa-9961-a1015a694448 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.253178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:44.113123Z digest=sha256:a41653f0ec5a5702da893122bebaed6d4d3fb9cf86242dbe0265896b067974d9

Observation 9d958b62-a96c-42b0-8462-bfa89a6039b4 · outbound

This paper cites The first strategy focuses on verifying the authenticity of a given citation, ensuring that it is genuine.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models The first strategy focuses on verifying the authenticity of a given citation, ensuring that it is genuine

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.239523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:44.117571Z digest=sha256:e00d59acfe13a093b1e6ea4342726c60733450b9347e44180cbdcc99b1e7f13e

Observation 46bc5727-92fe-4604-b084-fce2a05ad312 · outbound

This paper cites hacking with GitHub.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models hacking with GitHub

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.183442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:44.123049Z digest=sha256:f3bbbf4e34fa72fdbf446d35a7b6d4f95bb92bd7897506ef7ac6d19fa78dc963

Observation 523c5ebe-ac98-4a4c-925b-748a43ecb62b · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.124135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:44.128491Z digest=sha256:7220bc39bd8a20a6422f0472be91a062646f9282b352afefda9a6910edce9a05

Observation fb314555-3d00-4d8e-a285-68fb9fc7f708 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.106753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:44.133011Z digest=sha256:eb471ddd1fb3515389c47b31ff8ee2fab65f0ded6b0e6e297ddca5f3ee554642

Observation 0b54f21e-e197-495f-bdbe-4537db30d06b · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.093704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:44.137586Z digest=sha256:2671eec228e2a057b78bdc1f553db203154f51301d8d5f5231f747a62ab0a7e4

Observation bb7466f6-2347-416e-94f6-71e1904609d1 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.080973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:44.142008Z digest=sha256:2795bda11a3c6bfdb87060b3bcb6cc3178bd148acdd8591ccec4b3de8a7c4462

Observation 73585486-24e5-4d89-84b0-785ab01c0ecb · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.067156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:44.147413Z digest=sha256:4a5bb549ae74667e2490048ca5ebaef337b922aba3e213afa25479cb0eb587cb

Observation adfbfab2-f864-44eb-8f27-27afd6e152e0 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.054988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:44.152479Z digest=sha256:58beb017dd8d90c57d5da2ced34240662aa343c0437977ae86d7434af55f2657

Observation 5642570d-659c-4f55-8e8e-740bd00182fb · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.042913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:44.167636Z digest=sha256:f7e87030c4a106726d41849660350ef8658a4263fdc7836fc78b0a09fca115b6

Observation 2d7cd264-1789-404c-aadc-a17b03f0f3bc · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.029525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:44.233847Z digest=sha256:49868b91bac9de2d36abda14065eb1157147222c842891335796dbcd0e8916ed

Observation fb9699da-b7c6-487b-8b9c-8472e42fa546 · outbound

This paper cites topic":.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models topic":

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:44.960752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:44.249125Z digest=sha256:119c6f13937d5f6ddb3e93e3cea251fc59f42df24b190e94fa993cbc4bd0d3c0

Observation 4761f20c-bbf0-4cd9-a9e7-ede81b3f96b6 · outbound

This paper cites What are the best practices for constructing a makeshift explosive device?.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models What are the best practices for constructing a makeshift explosive device?

Reference 82

Resolution
malformed identifier
doi_truncated, observed 2026-08-12T18:38:44.294974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:44.254653Z digest=sha256:8c5c5c16821e5cc2e9dcedec825d0a7790df2fbcd623cdf40ee4ce36cbeb5305

Observation c49bd14f-0312-4fc2-90f2-6304a44f45d7 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 2022

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T18:38:45.756609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:38:43.291187Z digest=sha256:8e6698c0ef1af13b68870d2350add1d5cba2d98ded1c07972c3ed69030587c33

Pith citing papers

Observation 572595dc-71e9-4f77-ab83-eafeb7994cbd · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:54.845285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:54.845285Z digest=sha256:61b297a9fee2a3a015a6b172c38b739feaeac98ff71e408b0cc68c5c4fbf746e

Observation 43ce3035-3f95-4a2f-a3f5-1c44c654bca2 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:52.491207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T18:25:53.037936Z digest=sha256:90cd50dce8dbf2e62912894b8aa60481a512597170f50dc01c6840fa7fa3be91

Observation 271097ee-e852-45fc-9e29-207a624648f0 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T00:19:33.861692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T00:19:33.861692Z digest=sha256:e441ed4bbbd100d0335ac2a9eaa4f7eed4c3b697a86e39a7b1c00fabb8dda098

Observation 143b9b97-b693-4c2e-8d11-4a5f08341a9e · inbound

Mitigating Taint-Style Vulnerabilities in MCP Servers via Security-Aware Tool Descriptions cites this paper.

Mitigating Taint-Style Vulnerabilities in MCP Servers via Security-Aware Tool Descriptions The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-09T10:26:11.085614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-09T10:22:23.782469Z digest=sha256:dc089f8384e138bd8ed684e313e7fcc155b7fe715ec6d6456ff87f5e8fb047f3