Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T18:38:44.254653Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 4 inbound Pith citation observations for arXiv:2411.11407.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T18:38:44.254653Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:54.845285Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T10:26:11.084336Z
79 of 79 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ea7e5a30-1158-4eec-b4bd-005b8d4f8de7 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f9886c2-c391-4bf9-815d-b694e5b1d601 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 202891bd-cca9-421d-b686-666a4b8d8182 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Yampolskiy
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2e73b9a7-1067-4db5-aff2-abf7f21efea5 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Fundamental Limitations of Alignment in Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9be3bc26-89b2-462c-8f13-87b817b61c60 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”From Chatbots to Phishbots?: Phishing Scam Generation in Commercial Large Language Models.” 2024 IEEE Symposium on Security and Privacy (SP)
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 59e9f0c1-9e59-450b-be66-c901d62c7f65 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Sneakyprompt: Jailbreaking text-to-image gen- erative models.” 2024 IEEE symposium on security and privacy (SP)
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0f9959bd-ecc2-4f87-b674-789dd1e8eb43 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Superintelligence: Paths, dangers, strategies.” (2016): 196-203
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f3525391-670d-49c7-9139-2a46b0bda2cb · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Concrete Problems in AI Safety
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4436b0c3-136f-4b3b-b137-dd4df826334f · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Training language models to follow instructions with human feedback.” Advances in neural information processing systems 35 (2022): 27730-27744
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d9729422-dbb3-4cb6-9bbd-c8d274a0cb23 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Constitutional AI: Harmlessness from AI Feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 224fc152-3be0-4388-8e5b-1cb20ae7f16f · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb5f5474-da0f-45a1-a2c7-0a73a7319a82 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”A survey of reinforcement learning from human feedback.” arXiv preprint arXiv:2312.14925 (2023)
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18c42d46-be82-4d21-9467-9a5afea392c9 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1157d79f-03f3-41dd-adff-36e3b7fc7616 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.” arXiv preprint arXiv:2407.01599 (2024)
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db5ac0c2-eccb-4715-939d-720f960d7571 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ab34d54-d63e-4994-94e5-354755c9ae1b · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”On large language models’ resilience to coercive interrogation.” 2024 IEEE Symposium on Security and Privacy (SP)
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2b807fc4-8153-45bb-8bdc-54266768d9c9 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Multitask Prompted Training Enables Zero-Shot Task Generalization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dbeabe1-3bf4-4ea0-a973-ac5af287d11e · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Large Language Models: A Survey
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a8c8464-9397-4590-8b2f-ebe09aac644a · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Datasets for Large Language Models: A Comprehensive Survey
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9209f102-bce5-41d5-b6e0-dedf3352d1ea · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01f10d39-602a-4935-9a2a-100e06113f02 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models GPT-4 Technical Report
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22e378e5-ebf9-46a8-a3f4-1679e6424700 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f09a4bb1-f1cf-4f94-ac82-f1f7f84c5fad · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”The Claude 3 Model Family: Opus, Sonnet, Haiku.” Semantic Scholar, Corpus ID: 270640496 (2024)
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 904e9013-bc16-4d9a-8a79-3629d8f55316 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models On the Opportunities and Risks of Foundation Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 981db3f1-20d9-40c6-95d7-370815c32daa · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Extracting training data from large language models.” 30th USENIX Security Symposium (USENIX Security 21)
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0fb40c3c-b7b3-4e0c-9ed7-ee4933e58ccf · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Ethical and social risks of harm from Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2709f90-3b39-431c-9932-a7225d31aeb9 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Introducing ChatGPT
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5fdc6626-2ef2-422a-ba5e-04b90e9323bb · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Introducing Gemini: our largest and most capable AI model
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 47413d91-d54e-4c80-9b40-3fe436cc813a · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Introducing Claude
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cf5f9be5-0e35-40c1-88f5-99e6833f6d63 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Introducing the new Bing: The AI-powered assistant for your search
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 81b63bbe-3def-4bcb-8c5d-e368e5ec89cf · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0688ff11-6205-4a65-a96e-6db5894bf10d · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”LLM-Mod: Can Large Language Models Assist Content Moderation?.” Extended Abstracts of the CHI Conference on Human Factors in Computing Systems
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7480b75a-5aef-4ee8-8cc5-ad674c804770 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ShieldGemma: Generative AI Content Moderation Based on Gemma
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dee165cb-9a88-422f-a1da-c10f0725bcb3 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 116905c3-65b8-486b-8be9-7045ea8274e6 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9acb8d1b-1d1e-4d39-b627-cc6318e2be33 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”A new generation of perspective api: Efficient multilingual character-level transformers.” Proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6086b163-db9b-483f-b5b5-7a21ce015eab · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Using GPT-4 for content moderation.” OpenAI Blog (2023)
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8710398f-4350-4c40-bab8-0a6dfe397d54 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Influence of External Information on Large Language Models Mirrors Social Cognitive Patterns
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2cf65bf-ffe3-4768-ae5a-d790dd3cce55 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfc705b3-74ff-4fa3-a684-0186f148d364 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models LoFT: Local Proxy Fine-tuning For Improving Transferability Of Adversarial Attacks Against Large Language Model
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1617e55-54c5-4d39-a8f7-0d2af286fecf · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8c1be4b-96de-444b-a835-d9dfadceae12 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Judging llm-as-a-judge with mt-bench and chatbot arena.” Advances in Neural Information Processing Systems 36 (2023): 46595-46623
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7a49f01d-f323-4f8a-99f1-f430db9d7f44 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models DeepInception: Hypnotize Large Language Model to Be Jailbreaker
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcd8b79d-3989-4df2-869c-fe103429abe3 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aa00ae4-08ff-4f05-816f-31c4ab22223f · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70d2f0de-75b2-4ea6-ba45-9b73e5824165 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcab0098-51e3-4190-ad85-a22768b164b5 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2de9aca9-ce17-49cf-9ba2-05ffe3a16557 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c84d659d-372c-4bad-b3db-ba9ab52c269b · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Moderation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 09a69021-b621-4436-84bb-a28689777d7d · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models GPT-4o System Card
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c67792c0-89c7-46fe-a62a-0b0f42431819 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4efb6df-8a59-471c-a4dd-a7ffde26d3be · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7ffed2a-a3a6-40f1-b947-b3953e6135d3 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c954ad1c-002c-46b5-84ea-d44ef6beffec · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”R ´enyi divergence and Kullback-Leibler divergence.” IEEE Transactions on Information The- ory 60.7 (2014): 3797-3820
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dc06f73f-28f8-44e8-a5c5-62ee76df3e2a · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5545a5bd-3462-468d-87b5-e81da0ea1d58 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Baichuan 2: Open Large-scale Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13368495-a211-4134-80ee-a4ee4fa9d66a · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c5b9e95-3bbb-4406-ad7b-4cafa8674c50 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eaa7a6d-a457-4fce-a8b2-aeb24cca5847 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 307d9157-6827-46a3-9610-336acdc3a7a5 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65e53ffb-e955-4604-b319-5c62f8048be5 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Retrieval-augmented generation for knowledge- intensive nlp tasks.” Advances in Neural Information Processing Sys- tems 33 (2020): 9459-9474
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ae450bc7-ef8e-47d6-aa0a-2f4674a1b9b1 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Retrieval-Augmented Generation for Large Language Models: A Survey
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e7867e1-076f-460a-a517-32f07aa668cc · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Many-shot jailbreaking.” Anthropic, April (2024)
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cd0f20d4-aa5e-4a01-a080-df0a8cd022c1 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Attention is all you need.” Advances in Neural Infor- mation Processing Systems (2017)
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b96c4d92-fe81-41c2-a8c7-e09d36bc5aa8 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9e5eb1e2-1e4d-4baa-9961-a1015a694448 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9d958b62-a96c-42b0-8462-bfa89a6039b4 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models The first strategy focuses on verifying the authenticity of a given citation, ensuring that it is genuine
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 46bc5727-92fe-4604-b084-fce2a05ad312 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models hacking with GitHub
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 523c5ebe-ac98-4a4c-925b-748a43ecb62b · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fb314555-3d00-4d8e-a285-68fb9fc7f708 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0b54f21e-e197-495f-bdbe-4537db30d06b · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bb7466f6-2347-416e-94f6-71e1904609d1 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 73585486-24e5-4d89-84b0-785ab01c0ecb · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation adfbfab2-f864-44eb-8f27-27afd6e152e0 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5642570d-659c-4f55-8e8e-740bd00182fb · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2d7cd264-1789-404c-aadc-a17b03f0f3bc · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fb9699da-b7c6-487b-8b9c-8472e42fa546 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models topic":
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4761f20c-bbf0-4cd9-a9e7-ede81b3f96b6 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models What are the best practices for constructing a makeshift explosive device?
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c49bd14f-0312-4fc2-90f2-6304a44f45d7 · outbound
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 572595dc-71e9-4f77-ab83-eafeb7994cbd · inbound
SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43ce3035-3f95-4a2f-a3f5-1c44c654bca2 · inbound
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 271097ee-e852-45fc-9e29-207a624648f0 · inbound
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 143b9b97-b693-4c2e-8d11-4a5f08341a9e · inbound
Mitigating Taint-Style Vulnerabilities in MCP Servers via Security-Aware Tool Descriptions The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.