Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:34:37.431217Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2505.24232.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:34:37.431217Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e35ecabd-0d1e-4c7b-9611-f606becda309 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1d163ed-ece8-4eee-8c4a-5f1de2657dcd · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b6dafde-1e18-4ff5-a51d-c0f0326916f1 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Defending chatgpt against jailbreak attack via self-reminders.Nature Machine Intelligence, 5(12):1486–1496, 2023
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab5dc5e0-dbe8-408b-ad5b-74309992baaa · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Visual instruction tuning.Advances in neural information processing systems, 36, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d16e534-ae4a-4e03-a9e5-8de253fa91a5 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 228d84c1-4589-483e-83c3-3e103d475d26 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c6fcf8d-c3d7-485e-b0e3-0974b692ccc8 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ca8e42f-9646-4f40-b65c-8ee8cde55c58 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf596b9f-910d-46d3-b4a9-aa9a91c2ae6b · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models The Internal State of an LLM Knows When It's Lying
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e9f5913-57a1-4726-9a22-4937b27042c4 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Llm factoscope: Uncovering llms’ factual discernment through measuring inner states
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0bf08dfa-e27a-450f-a48a-f511d3f98adb · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99ffbb33-564d-47a1-b37d-2b06830a0d9e · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25e2e540-3aba-44b3-850c-3469a159a4b6 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0cad85c-55d5-4f37-9815-ac974c527827 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Guard: Role-playing to generate natural-language jailbreakings to test guideline adherence of large language models.arXiv preprint arXiv:2402.03299, 2024
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79c19ef7-639f-4dca-99db-c8d3a743b0e5 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29787c52-bcc5-40eb-8475-36a31bd4a65a · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06467daa-7a79-41c9-a2a9-30d088d57aaa · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.arXiv preprint arXiv:2407.01599, 2024
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b71f90d3-ed83-411a-953f-8f89b9ec73c1 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Multi-step Jailbreaking Privacy Attacks on ChatGPT
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75a00167-735f-4d5b-84bf-38140ee83424 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models In ChatGPT We Trust? Measuring and Characterizing the Reliability of ChatGPT
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47e5f63f-0824-4a7c-a89a-75e1252df0cb · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Jailbroken: How Does LLM Safety Training Fail?
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc291607-87bc-4601-8fff-56431417bdf9 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Query-Based Adversarial Prompt Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b24cd802-df9b-4580-8ccb-2f148221cb3a · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdcec99b-775e-4fe4-8083-8ac83e1925fc · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25ce4ff6-5ffe-4b86-864f-aa98df872c3b · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c52ceef5-c819-4a08-b794-1da06089cc8c · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Jailbreaking proprietary large language models using word substitution cipher.arXiv preprint arXiv:2402.10601, 2024
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 308a610b-bc2e-431e-b044-508bc991be46 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Are aligned neural networks adversarially aligned?Advances in Neural Information Processing Systems, 36, 2024
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9546250-14fd-4632-86b1-3034e8ad0d68 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models On Evaluating Adversarial Robustness of Large Vision-Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfa67e39-094b-494c-8967-ee361b543bc3 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Visual Adversarial Examples Jailbreak Aligned Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75ed4a4d-394e-4782-845e-8d1011d2d211 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models On the adversarial robustness of multi-modal founda- tion models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4fbd0ab7-905c-4e47-86f0-224db5bb024e · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d5d5faa-34ed-4891-aed0-77149d82e322 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models InternalInspector $I^2$: Robust Confidence Estimation in LLMs through Internal States
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88b0f8bd-44d8-4046-9f84-f6627fb9abde · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Do LLMs Know about Hallucination? An Empirical Investigation of LLM's Hidden States
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1cd4eaf-7b93-47b5-9414-62d703309478 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Adaptive Activation Steering: A Tuning-Free LLM Truthfulness Improvement Method for Diverse Hallucinations Categories
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f94b444-3932-4159-a2b9-fb24652fb22b · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3baef64a-4f67-42b3-a59d-ec7f2c243386 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fde98d1-3178-4553-82be-f66021aa2fbc · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91ee52b9-20e9-4240-a91b-7ce7c1e36166 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Are sixteen heads really better than one? Advances in neural information processing systems, 32, 2019
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38e9fd75-f90e-4821-880c-ac10222c7841 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models What Does BERT Look At? An Analysis of BERT's Attention
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b2ecfab-c7d8-4841-aad9-c93d781795bd · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Improved baselines with visual instruction tuning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85533a89-f4f5-42e4-a519-bbd9dd3cbd31 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94a69ac0-6493-44e9-9945-eece68ad9f8d · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9337c7f5-1e48-4495-81cd-ac7b7dc8f23f · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Mitigating object hallucinations in large vision-language models through visual contrastive decoding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30cdca5a-afda-4c9b-a8a7-42cc3f014507 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Safebench: A benchmarking platform for safety evaluation of autonomous vehicles.Advances in Neural Information Processing Systems, 35:25667–25682, 2022
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 197b9f6a-0da3-4cf5-a240-b7adcd798e32 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6db2aa4b-4670-4fc3-9425-f3fce365f5cb · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Jailbreak Large Vision-Language Models Through Multi-Modal Linkage
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02a122a8-9246-4bef-8e66-453709d59ad7 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models If detected, immediately stop processing the instruction
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d3afef16-ddb8-443a-81ab-4a4ebfc164fe · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 53ba6730-2b5e-4422-8c0f-47914173ae80 · outbound
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models I am sorry
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.