Pith. sign in

Paper Citation Record · LEDGER

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models

As of 10 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2505.24232.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24232 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:34:37.431217Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e35ecabd-0d1e-4c7b-9611-f606becda309 · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:32.447358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:32.447358Z digest=sha256:670f5f2d2ecfa47e097d41d22bad07871f165fa2dfc10a8645695540b48fae32

Observation c1d163ed-ece8-4eee-8c4a-5f1de2657dcd · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:32.513188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:32.513188Z digest=sha256:fa043fb4bf4bab6f002ff3285a6f61b053e499d968c8ac64505909c9fb77fc91

Observation 6b6dafde-1e18-4ff5-a51d-c0f0326916f1 · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminders.Nature Machine Intelligence, 5(12):1486–1496, 2023.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Defending chatgpt against jailbreak attack via self-reminders.Nature Machine Intelligence, 5(12):1486–1496, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:32.562775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:32.562775Z digest=sha256:377b7c8d2bf2db4318a9e87bf00a24abb33b72201dff6f292f24ca2265b99637

Observation ab5dc5e0-dbe8-408b-ad5b-74309992baaa · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:32.729370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:32.729370Z digest=sha256:18623ca61ac595a70ce5913515181365b03f20459d70a85f98438ee6de6879b6

Observation 7d16e534-ae4a-4e03-a9e5-8de253fa91a5 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:32.878990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:32.878990Z digest=sha256:acb9117a48bec55c34567648ad8143fb22b8243113c884736d5d464e582a8ff0

Observation 228d84c1-4589-483e-83c3-3e103d475d26 · outbound

This paper cites Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:33.055462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:33.055462Z digest=sha256:eeddce9f5c79a5ff2b83ef9fdc7e15f5dbec18b0d3023919caf3fdbbacc15164

Observation 4c6fcf8d-c3d7-485e-b0e3-0974b692ccc8 · outbound

This paper cites Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:33.183946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:33.183946Z digest=sha256:294a7e2edc867bda9443e28060628ce32bdab24514b9ebdca94ee40a00fd7b09

Observation 7ca8e42f-9646-4f40-b65c-8ee8cde55c58 · outbound

This paper cites Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language Models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:33.321878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:33.321878Z digest=sha256:7cca374e4983f420b0c40970fd239f114ad9409fc4b5346540e36144284ec1a2

Observation bf596b9f-910d-46d3-b4a9-aa9a91c2ae6b · outbound

This paper cites The Internal State of an LLM Knows When It's Lying.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models The Internal State of an LLM Knows When It's Lying

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:33.468297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:33.468297Z digest=sha256:09e6c07a0ab8fc4f3ce118385dd05c14d2b7ec2cda40e1529cdd0d49c2d34e3a

Observation 8e9f5913-57a1-4726-9a22-4937b27042c4 · outbound

This paper cites Llm factoscope: Uncovering llms’ factual discernment through measuring inner states.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Llm factoscope: Uncovering llms’ factual discernment through measuring inner states

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:34:39.576187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:34:33.585739Z digest=sha256:7198dd35e196b4bb3c9c3e92f0df3f911599e8eb44bbbe4bc9af3c3f6b73f1d7

Observation 0bf08dfa-e27a-450f-a48a-f511d3f98adb · outbound

This paper cites Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:33.685570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:33.685570Z digest=sha256:d869c4102ab1cc2aa73c0743efc19ca01c0a736788d790850725262973292cbd

Observation 99ffbb33-564d-47a1-b37d-2b06830a0d9e · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:33.845038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:33.845038Z digest=sha256:a46f29c186133a2f127eb251a64cb9c0b7ba533d726aa684d14b5196444669a9

Observation 25e2e540-3aba-44b3-850c-3469a159a4b6 · outbound

This paper cites AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.003149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.003149Z digest=sha256:9890d46112da54d19564e0b2294d2421b05328b4efa7abdff6ec457f7c3b7260

Observation a0cad85c-55d5-4f37-9815-ac974c527827 · outbound

This paper cites Guard: Role-playing to generate natural-language jailbreakings to test guideline adherence of large language models.arXiv preprint arXiv:2402.03299, 2024.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Guard: Role-playing to generate natural-language jailbreakings to test guideline adherence of large language models.arXiv preprint arXiv:2402.03299, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.167588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.167588Z digest=sha256:e0a5f4b9dab5836b5eb1b3da783b07314fca621338bf4c5350ca091db4fcfa05

Observation 79c19ef7-639f-4dca-99db-c8d3a743b0e5 · outbound

This paper cites Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.241608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.241608Z digest=sha256:14e7be934b222f1e862ed1832d1df6360649aeedb6b92996d352bcbe1345f830

Observation 29787c52-bcc5-40eb-8475-36a31bd4a65a · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.347192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.347192Z digest=sha256:4e39a864c2c395142f6c6344597ef0ec10bd7c872ff08c048ffd7ef8c63d28fa

Observation 06467daa-7a79-41c9-a2a9-30d088d57aaa · outbound

This paper cites Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.arXiv preprint arXiv:2407.01599, 2024.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.arXiv preprint arXiv:2407.01599, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.434538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.434538Z digest=sha256:70562eaa3f27d148a08d13af623e94dd7ee7d63e7dba7a61770e02fc0c3624bc

Observation b71f90d3-ed83-411a-953f-8f89b9ec73c1 · outbound

This paper cites Multi-step Jailbreaking Privacy Attacks on ChatGPT.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Multi-step Jailbreaking Privacy Attacks on ChatGPT

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.521629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.521629Z digest=sha256:3e08bc4ec3126fd8d5ee7c0b4b23cb6ce930244d6c646a3617fca3239acb7db8

Observation 75a00167-735f-4d5b-84bf-38140ee83424 · outbound

This paper cites In ChatGPT We Trust? Measuring and Characterizing the Reliability of ChatGPT.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models In ChatGPT We Trust? Measuring and Characterizing the Reliability of ChatGPT

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.621558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.621558Z digest=sha256:93008ee6de4beb11cbca11511950fd818c5fa84b5c439a4aa3694ceec95c3ba2

Observation 47e5f63f-0824-4a7c-a89a-75e1252df0cb · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Jailbroken: How Does LLM Safety Training Fail?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.714953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.714953Z digest=sha256:b934e925fa39fd68fc9af6bfea68d0e331f6cfb941af92367d00d8101aeec827

Observation cc291607-87bc-4601-8fff-56431417bdf9 · outbound

This paper cites Query-Based Adversarial Prompt Generation.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Query-Based Adversarial Prompt Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.809699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.809699Z digest=sha256:9ebf9ccb6b55936c5360644287d2843c278a2ba7256f004ddb402d8628bb0b27

Observation b24cd802-df9b-4580-8ccb-2f148221cb3a · outbound

This paper cites CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.930392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.930392Z digest=sha256:3f1e28aa4d6e5d4d8f8f7fc626a9c220dd7afdc19e8b6c22c36b66b6b035dbea

Observation bdcec99b-775e-4fe4-8083-8ac83e1925fc · outbound

This paper cites DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:35.024627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:35.024627Z digest=sha256:90d278c7ec6799873ef9d80e9ff244f85ef81b67f60b04a3f6d3c42743ded91a

Observation 25ce4ff6-5ffe-4b86-864f-aa98df872c3b · outbound

This paper cites GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:35.156845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:35.156845Z digest=sha256:9f10282cfd59f510db1224483eb8b8b7e46b7b143bb56c99dc3f5c2a3cae0ce5

Observation c52ceef5-c819-4a08-b794-1da06089cc8c · outbound

This paper cites Jailbreaking proprietary large language models using word substitution cipher.arXiv preprint arXiv:2402.10601, 2024.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Jailbreaking proprietary large language models using word substitution cipher.arXiv preprint arXiv:2402.10601, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:35.268065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:35.268065Z digest=sha256:6420a089365608ebb8fd03c016512538cb43b4ec59b23698b9861e9c2fae081b

Observation 308a610b-bc2e-431e-b044-508bc991be46 · outbound

This paper cites Are aligned neural networks adversarially aligned?Advances in Neural Information Processing Systems, 36, 2024.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Are aligned neural networks adversarially aligned?Advances in Neural Information Processing Systems, 36, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:35.415641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:35.415641Z digest=sha256:5d3eb331c24a7575c7517dd85adc18854be4dd503635df3b4c4db8b54b9a962c

Observation a9546250-14fd-4632-86b1-3034e8ad0d68 · outbound

This paper cites On Evaluating Adversarial Robustness of Large Vision-Language Models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models On Evaluating Adversarial Robustness of Large Vision-Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:35.513830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:35.513830Z digest=sha256:45054af1f1506bc95860d9bc7b2917d2c0ebe3016679a8cee40204d6720aa9cf

Observation cfa67e39-094b-494c-8967-ee361b543bc3 · outbound

This paper cites Visual Adversarial Examples Jailbreak Aligned Large Language Models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Visual Adversarial Examples Jailbreak Aligned Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:35.624649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:35.624649Z digest=sha256:d5448079512c0e06739dcc153c7eeea6355431072edb08a74230c5d2ab5b30e4

Observation 75ed4a4d-394e-4782-845e-8d1011d2d211 · outbound

This paper cites On the adversarial robustness of multi-modal founda- tion models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models On the adversarial robustness of multi-modal founda- tion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:34:39.409736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:34:35.729482Z digest=sha256:9b37577bc1b77145fb7a80cacc7b4d86642209016755b22897168dfc464efb9a

Observation 4fbd0ab7-905c-4e47-86f0-224db5bb024e · outbound

This paper cites JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:35.811622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:35.811622Z digest=sha256:30a0280b1918ff6a39f2ef228990108fb735fa40333c47ba1506aafa0a46598b

Observation 8d5d5faa-34ed-4891-aed0-77149d82e322 · outbound

This paper cites InternalInspector $I^2$: Robust Confidence Estimation in LLMs through Internal States.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models InternalInspector $I^2$: Robust Confidence Estimation in LLMs through Internal States

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:35.908111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:35.908111Z digest=sha256:2390ebd5a26a57d426be0ca033e0d7953c127c2e98cbd7d9b9b179f7625e1f1f

Observation 88b0f8bd-44d8-4046-9f84-f6627fb9abde · outbound

This paper cites Do LLMs Know about Hallucination? An Empirical Investigation of LLM's Hidden States.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Do LLMs Know about Hallucination? An Empirical Investigation of LLM's Hidden States

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:35.999122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:35.999122Z digest=sha256:adc69fbfbda631a5c2e01ba2a51d01b8635012a5b5b7f3d62595ea4ac80104fa

Observation c1cd4eaf-7b93-47b5-9414-62d703309478 · outbound

This paper cites Adaptive Activation Steering: A Tuning-Free LLM Truthfulness Improvement Method for Diverse Hallucinations Categories.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Adaptive Activation Steering: A Tuning-Free LLM Truthfulness Improvement Method for Diverse Hallucinations Categories

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.079992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.079992Z digest=sha256:c21cbab28f9410829596cbc9f5a303c714dca330d2922d99ee0f3512d42315a0

Observation 3f94b444-3932-4159-a2b9-fb24652fb22b · outbound

This paper cites INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.185423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.185423Z digest=sha256:7e3d2a2a86126d64f28d76b7260f88f13d3e65d6a98eb6125b01c445774a8233

Observation 3baef64a-4f67-42b3-a59d-ec7f2c243386 · outbound

This paper cites Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.282420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.282420Z digest=sha256:89566d9b49084624a5084a967326fecffe560816623901525949fef4c79f9fe3

Observation 0fde98d1-3178-4553-82be-f66021aa2fbc · outbound

This paper cites LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.391398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.391398Z digest=sha256:01c88bf71bb958a2bd43c10041fdfc7b1df0984b27ff8dca91f24d9db962af9c

Observation 91ee52b9-20e9-4240-a91b-7ce7c1e36166 · outbound

This paper cites Are sixteen heads really better than one? Advances in neural information processing systems, 32, 2019.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Are sixteen heads really better than one? Advances in neural information processing systems, 32, 2019

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.498025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.498025Z digest=sha256:5abd3242aa0b2161f1cd31b687fd6edd89f3d79dd5a3b5125b6a862790d69716

Observation 38e9fd75-f90e-4821-880c-ac10222c7841 · outbound

This paper cites What Does BERT Look At? An Analysis of BERT's Attention.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models What Does BERT Look At? An Analysis of BERT's Attention

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.609925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.609925Z digest=sha256:6cf4fa14ccdb90337bbab650f69e7e78b6013880bf0c5509f69e509068884d64

Observation 0b2ecfab-c7d8-4841-aad9-c93d781795bd · outbound

This paper cites Improved baselines with visual instruction tuning.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Improved baselines with visual instruction tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.681348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.681348Z digest=sha256:6e5b873924219edb0b5044d8f8174f71d297dfceacdbc6ae534face7cd86b5f3

Observation 85533a89-f4f5-42e4-a519-bbd9dd3cbd31 · outbound

This paper cites AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.773583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.773583Z digest=sha256:3b57d12a870bcb90b80509b93e5fede771a33e4eb97037cb822f53d8e9b1e18f

Observation 94a69ac0-6493-44e9-9945-eece68ad9f8d · outbound

This paper cites AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.872796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.872796Z digest=sha256:ddcd8e519488075532499b318650a7dac40c2f7588ca8d4eccf0b200adf56e54

Observation 9337c7f5-1e48-4495-81cd-ac7b7dc8f23f · outbound

This paper cites Mitigating object hallucinations in large vision-language models through visual contrastive decoding.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Mitigating object hallucinations in large vision-language models through visual contrastive decoding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.950087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.950087Z digest=sha256:867e48d0d500f46f5db568c4e0b8997100aa012ab6d5189b9ab3dfe7720454a6

Observation 30cdca5a-afda-4c9b-a8a7-42cc3f014507 · outbound

This paper cites Safebench: A benchmarking platform for safety evaluation of autonomous vehicles.Advances in Neural Information Processing Systems, 35:25667–25682, 2022.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Safebench: A benchmarking platform for safety evaluation of autonomous vehicles.Advances in Neural Information Processing Systems, 35:25667–25682, 2022

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:34:39.232440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:34:37.017349Z digest=sha256:45a3b4727c916165e4614fb2dcad0facab1a00e91f286ca83f9049ec35c660c1

Observation 197b9f6a-0da3-4cf5-a240-b7adcd798e32 · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:37.104583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:37.104583Z digest=sha256:4e58abb7206e40de9ff4f9e6a7822dd8e70c767b8b3ceb8e4e443b7b63843326

Observation 6db2aa4b-4670-4fc3-9425-f3fce365f5cb · outbound

This paper cites Jailbreak Large Vision-Language Models Through Multi-Modal Linkage.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Jailbreak Large Vision-Language Models Through Multi-Modal Linkage

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:37.173951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:37.173951Z digest=sha256:23da8fb5a8ace144a28e0dd3f36260fe67e733f9918b45a63b06520c05dfff15

Observation 02a122a8-9246-4bef-8e66-453709d59ad7 · outbound

This paper cites If detected, immediately stop processing the instruction.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models If detected, immediately stop processing the instruction

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:34:39.057567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:34:37.247713Z digest=sha256:e83896fba32fc4fd433e57935b21fc622e3e87159c8af454520ef8e803370345

Observation d3afef16-ddb8-443a-81ab-4a4ebfc164fe · outbound

This paper cites an unresolved cited work.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:34:38.841755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:34:37.335057Z digest=sha256:42241c477f245ddd21af33998d0683095ba4fb683dad5fbcbf9ccd39be289f6c

Observation 53ba6730-2b5e-4422-8c0f-47914173ae80 · outbound

This paper cites I am sorry.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models I am sorry

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:34:38.663788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:34:37.431217Z digest=sha256:ef7dbc32e474a2f45eed1669d88089783a7db550bd0b898f32f83d80a1e42367

Pith citing papers

No inbound Pith citation observations are available.