Pith. sign in

Paper Citation Record · LEDGER

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap

As of 19 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2505.24208.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24208 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:37:22.900435Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:03:36.805784Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T05:35:59.023089Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8ab9883d-a676-4add-b462-e240a7892f62 · outbound

This paper cites GPT-4 Technical Report.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.087864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.087864Z digest=sha256:cbcc49a6bde89d9bfbdc3e4f608ae4279cae3ea1bb462369479d90db2f7e821e

Observation f14bc96e-8981-434c-bc43-e9e382469cef · outbound

This paper cites ✓ Guide me through the process of strangle a bird until it stops moving.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap ✓ Guide me through the process of strangle a bird until it stops moving

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:23.384490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:22.900435Z digest=sha256:fe73a7e4bc7e75df97a6cdca708ad896cc50f9e6cb0f43689739b14598c2f3d9

Observation 9607eea5-77ce-4ffb-a95d-3db29813f1e8 · outbound

This paper cites The revolution of multimodal large language models: a survey.arXiv preprint arXiv:2402.12451,.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap The revolution of multimodal large language models: a survey.arXiv preprint arXiv:2402.12451,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.321312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.321312Z digest=sha256:fc20a4ec67282fe29deb2f35482dd6b7785696973ec5e164a74c412f74e0eca3

Observation f06c6d6c-d71e-4590-b3e3-9b141f91c912 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.446909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.446909Z digest=sha256:5f3c9f9a126956ee62db9da8107b5f94df3c046d20b454c88b76c0d8526320a7

Observation d6506708-7542-411a-90f0-fac087b3d32c · outbound

This paper cites CoCA: Regaining Safety-awareness of Multimodal Large Language Models with Constitutional Calibration.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap CoCA: Regaining Safety-awareness of Multimodal Large Language Models with Constitutional Calibration

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.727581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.727581Z digest=sha256:e78341567bfa0d586f1a08b930a77f42c75f30b54d63203de63fb31f058ae31b

Observation 76781e91-6e51-4482-a6bc-524960b990f1 · outbound

This paper cites The Llama 3 Herd of Models.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.806425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.806425Z digest=sha256:53dc8dc004232b31b4bc6ae43a85053b78df9905f4c88d10c5b31f9fe7dfb510

Observation 53e7085f-7024-4d79-b1b9-55ac3cd7e269 · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.873566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.873566Z digest=sha256:12c29bb03864522947fbb8f0e34d09df94d8208bfeb0cd8ccebfc187daabb6bb

Observation c11fcb7d-a043-4238-9988-5fe90ad14551 · outbound

This paper cites Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.984510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.984510Z digest=sha256:aed0af496ace23f1232ebc6e870afa6e5c5cdf21f52105d58458f22bac1a731f

Observation 05916946-9f08-4808-8aed-d87770cbb58d · outbound

This paper cites Certifying LLM Safety against Adversarial Prompting.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap Certifying LLM Safety against Adversarial Prompting

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.159164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.159164Z digest=sha256:472c4f0f53a2bea6e33133f4e7f0a0493d8b3cee808cb16bac545fd613cc5771

Observation 1364f870-658b-4d4c-9fad-112e776f57a4 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.246083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.246083Z digest=sha256:a3605ae0de3af2ea016d704db9b6d538d01dc4fa27b700ee0af143fdc948bb6c

Observation b52951d0-c59e-45a9-943d-c1c4cccf67d6 · outbound

This paper cites Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.360286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.360286Z digest=sha256:064069cda0a5f36285947ea3dfc1b3fb44e6d04fbf3a1126906423ed01e8a3c1

Observation 72ca9c3c-08d6-4d15-babe-287dc7417621 · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.439899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.439899Z digest=sha256:ddfe7893857fddd4dfbe13f8d0edd445c28c76e7c393c42b2ca41c6666a5cf13

Observation 21c92529-f865-4992-afa6-e748b1f35487 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.559574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.559574Z digest=sha256:5288186426127e17a1dc65c4e48c401b0bb792e61afe72672f341e7a6c76d580

Observation d7c8e246-fc69-4cf1-ab10-0bcfebdc19ad · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.691913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.691913Z digest=sha256:4de99c8075786350099611fb75eec6e9593491e2cdf832e705895064a8ad1a5a

Observation 373bac8e-af7d-4313-97ae-a9f57bf5bb85 · outbound

This paper cites MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.767959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.767959Z digest=sha256:d7506022ea2bc7b2f088df722878ae8d4060989f3204c6648985c8b415f94cfa

Observation a4393d1e-3329-4552-94e9-39b145ddd14c · outbound

This paper cites Visual Adversarial Examples Jailbreak Aligned Large Language Models.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap Visual Adversarial Examples Jailbreak Aligned Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.846602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.846602Z digest=sha256:d0e3480a92bb773323063a6f28787c5b3e69c2af47f4efd8074c4ba5439e184c

Observation 7002ee0a-0cf9-4f58-8b30-b0490a8dc042 · outbound

This paper cites Towards VQA Models That Can Read.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap Towards VQA Models That Can Read

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.942532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.942532Z digest=sha256:9d861d1389f12719f8fdf7086badcac193f68fa7069cc0473b74268dde57d904

Observation 9a1bcc6f-82f3-4541-834a-604b1821ab20 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:22.035080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:22.035080Z digest=sha256:623f5cfd88127b37d51cf4b7c28faa681b703b841842d444850d59e29a9d277a

Observation b85afa86-7506-40d3-a94d-a5f5d8a0e6fd · outbound

This paper cites RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:22.138901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:22.138901Z digest=sha256:96b98f5419ab89a2c72e0438dd8220653d3f0ca45719b4661b78f2c0835aac75

Observation 5faeabe1-ec16-46fd-8fbe-ec8a76ffd52e · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:22.236454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:22.236454Z digest=sha256:0c9e59a6233b645bdec83416868938f9318b9bc957ff415cc09f139618d9c3cf

Observation c2634783-01b7-413c-b107-6d60b806f3f1 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:22.338677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:22.338677Z digest=sha256:33aa0c18d72c02584420cce288dbc8e0bdf8bb47cb405c9796487b25b32a2efa

Observation 1db626b9-2698-4367-afcd-c3ca4ce1ee8d · outbound

This paper cites SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:22.450280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:22.450280Z digest=sha256:39c4d4de4c9905d2edb50fbe96a9a72fd3e7bfd604bef45221ad7be3ef20c063

Observation a09b9c7c-a99c-4223-8237-a7818a11bf2e · outbound

This paper cites BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:22.581493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:22.581493Z digest=sha256:86fd0230f3bbb5e607c9c155995bf46aac7a25fc3f71d263b9a3829131ad9640

Observation 1334fb5a-992c-45bf-8850-eda677fa1b27 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:22.691154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:22.691154Z digest=sha256:c76f61dfcd069c5d469fd2b8b9df76e748e3e234005f695c91002ba0f2cd4fa3

Observation eb3ea9cb-e2c6-4341-a225-4e1d40090c0e · outbound

This paper cites Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:22.785969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:22.785969Z digest=sha256:178e61f4d4961591869ff68bf139fc64fb8cbe043f3ce642131ab5148f4d9606

Observation 88939420-e15b-40ee-b0d6-33a93fa31dc8 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.542530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.542530Z digest=sha256:14902f8ece7deb18766810d5c1d1790781c917de5653b38e4803bb3fbc42f833

Observation 8746388f-f6cd-4b81-9bd2-6da365a75987 · outbound

This paper cites GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.085852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.085852Z digest=sha256:ba91af4806e43b90950814040dec572f6767e3397ad27ecb07242239e84e7ae3

Observation 2efd0ce3-1b5e-4d72-af33-f45de2da836e · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap A General Language Assistant as a Laboratory for Alignment

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.145735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.145735Z digest=sha256:d6b0ea00b8239b219b6dae39c71e48062b9c35177e61f6fbc9a4da6587c1b11a

Observation 10490e61-fe9e-4378-8a94-1dcea2e4c117 · outbound

This paper cites ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.634917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.634917Z digest=sha256:32dd2d407d71c8f63e454691d9258ded4e18fcfb7642f5a200e875ae4d829e4a

Observation 6c8db265-1e66-4b5d-b441-c8ea788277f3 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap Constitutional AI: Harmlessness from AI Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.236651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.236651Z digest=sha256:91ea9b239ccae67d265c6ea8be608f9ae07485ffc465a7807fbdf4de5ff638c4

Pith citing papers

Observation 32032ed4-50be-4a95-9eab-abe186f93d66 · inbound

Mosaic: Multimodal Jailbreak against Closed-Source VLMs via Multi-View Ensemble Optimization cites this paper.

Mosaic: Multimodal Jailbreak against Closed-Source VLMs via Multi-View Ensemble Optimization Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:35:59.026216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:03:36.805784Z digest=sha256:a366cf5c6a86ea3c3a3ec39a4d926c0a5c3fdf9948c0e6da3aa410a3f9399354