Pith. sign in

Paper Citation Record · LEDGER

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities

As of 9 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2507.11155.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11155 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:21:35.386601Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact1
  • verified fuzzy32
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de479335-37c3-4d7c-be41-782cef9fa42e · outbound

This paper cites https://huggingface.co/deepseek-ai/ 13 DeepSeek-R1.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities https://huggingface.co/deepseek-ai/ 13 DeepSeek-R1

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:36.236155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.122949Z digest=sha256:91f94bd2e2046220fe89aadb6ce0d2351b4822228b5bad2f9a6448b59c0e0ad5

Observation 3845f66e-d594-4ce7-b367-fd4bae1e3fcd · outbound

This paper cites https://openai.com/index/gpt-4-1/.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities https://openai.com/index/gpt-4-1/

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:36.213411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.127530Z digest=sha256:1d47dadadd5aef8c1e78a1e38c3fc0872cb7714e46799f9577c5c9c0721402d3

Observation 2705751e-6dbd-4d2d-bf6d-25123558ba53 · outbound

This paper cites https://openai.com/research/gpt-4v- system-card.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities https://openai.com/research/gpt-4v- system-card

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:36.185507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.132351Z digest=sha256:d920f6d32d4f576b8dd3d02a3b26c4aa718c7e1cf1f7b8c29579cb279058da6b

Observation 62008d21-bb05-4e6a-b5c5-b5a8b18aa477 · outbound

This paper cites https://huggingface.co/datasets/laion/ laion2B-en.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities https://huggingface.co/datasets/laion/ laion2B-en

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:36.160262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.136636Z digest=sha256:d1b9c20649c2ad8eb7d5e35398251fb336be08173173fe644f994d2be3f5129e

Observation dfb1bc55-ccb3-4378-a25f-0c5bbf95b529 · outbound

This paper cites https: //academictorrents.com/details/ 1cda9427784a6b77809f657e772814dc766b69f5.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities https: //academictorrents.com/details/ 1cda9427784a6b77809f657e772814dc766b69f5

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:36.132459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.140805Z digest=sha256:addd631bb2560bef27121c75e150a7c556620c3a33135741a584e9456998e571

Observation 6ac9360b-cd54-4c35-bb10-ff6067519a5a · outbound

This paper cites https://web.archive.org/web/ 20220406151527/https://labs.openai.com/policies/ content-policy.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities https://web.archive.org/web/ 20220406151527/https://labs.openai.com/policies/ content-policy

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:36.106730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.145036Z digest=sha256:1f0e362efad1d6165f8ce5070ef221075bd18f91c85fc819675c322ba54659f6

Observation 83110e28-ebbf-498a-82ea-25f343e7677a · outbound

This paper cites https://openai.com/o1/.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities https://openai.com/o1/

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:36.085062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.148841Z digest=sha256:f4cdc48cf7a22d3b6396c4de1a6822335ba2c70d2b62ff68b3fc5e09fa835ffd

Observation e450028e-a585-4add-9526-52bf0d21a505 · outbound

This paper cites https://osf.io/2rqad/.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities https://osf.io/2rqad/

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:36.057871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.152404Z digest=sha256:8c9f2c2d1c01358c29d39376ea7b4d25fae69cd3c30a2de59a219de43331c667

Observation 18bdc356-6edd-44ff-8389-a227cc205a27 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.156208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.156208Z digest=sha256:2a77f5c374e5c627587a405a157cd3e710362a4fa725df2294d9f9083760154d

Observation 47a09fa3-174c-42c8-ba3d-c067bf6da595 · outbound

This paper cites Designing Neural Network Architectures using Rein- forcement Learning.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Designing Neural Network Architectures using Rein- forcement Learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:36.036711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.159839Z digest=sha256:b2ba71dd8959daf28746c8e3177a6ae796968bde4de1a0a6001bb0e58fb1568c

Observation 2c0199ff-0443-4701-9516-ef4c6ce22a61 · outbound

This paper cites Image Safeguarding: Reasoning with Conditional Vision Language Model and Obfuscating Unsafe Content Counterfactually.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Image Safeguarding: Reasoning with Conditional Vision Language Model and Obfuscating Unsafe Content Counterfactually

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:36.019184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.163361Z digest=sha256:2fa020902390fe91c8013920d6434ddd3f66c504d46138c1d38080a74c21c45b

Observation f4895b71-2732-4461-b5bc-b051b20bc177 · outbound

This paper cites InternLM2 Technical Report.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities InternLM2 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.166750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.166750Z digest=sha256:64e837dce2e33b6719284ee83dedd0272ae9d24eb9c2b296877a313b20b190fe

Observation abe12a9e-b13f-4f56-bd0c-71d1e175db75 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.170826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.170826Z digest=sha256:55126df7bb2aefbc3fcd62f3e6567b2d5ed5e4e31713dd25c24ee9a2fe63826f

Observation fd2e3fcc-0580-4157-94c0-15377cc20100 · outbound

This paper cites Christiano, Jan Leike, Tom B.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Christiano, Jan Leike, Tom B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.175318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.175318Z digest=sha256:2b0d128c40482de2586ac88709ce2f2e70335894c252fcb9945bc73a94786a19

Observation 86f736ee-9392-4189-ae26-0823d40af63c · outbound

This paper cites an unresolved cited work.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:21:35.997267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.179278Z digest=sha256:3b427f52bc3ffd46234513d324bd20243ec4ad0488da2ed9a6ac7e8d751ba9fc

Observation 3e83941c-494d-48dd-8b6b-bae0b4bfcd27 · outbound

This paper cites ImageNet: A large-scale hierarchical image database.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities ImageNet: A large-scale hierarchical image database

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.984232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.183752Z digest=sha256:7597aa961516a623a22c47e52e301eb08e0058847d01c0ec4faeac2165719912

Observation 46b7d000-77e8-41a7-a883-caccc8755ba0 · outbound

This paper cites ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.186888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.186888Z digest=sha256:e8bc93b38fe34e677664b06e46891f1ad4d0f2186ba0cf7e1b2d64243329aeeb

Observation c217dcee-af51-45ad-ac57-61bb58e5735b · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.190481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.190481Z digest=sha256:5f3e3c78debbc46e32d5c997265a20b9ffda2016c568d5cbaf0821a0f82cb249

Observation 1936df79-e57c-42bc-a82a-46f309452a11 · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.194359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.194359Z digest=sha256:7d1cb1f2892224fba336019fa70df648d781d2d14941ac7832eb2ffd51defd46

Observation 23b65e5f-de48-450a-883e-a0c63e09addf · outbound

This paper cites Fleiss’ kappa statistic without paradoxes.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Fleiss’ kappa statistic without paradoxes

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.198149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.198149Z digest=sha256:495c4673eb3b86bb49fb1accd1944e8c2444f1274cef82553a8e6e4539497d29

Observation 7de8e157-ab5d-4a0c-aaf3-bb59dcb0aaa9 · outbound

This paper cites Measuring Nominal Scale Agreement Among Many Raters.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Measuring Nominal Scale Agreement Among Many Raters

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.202068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.202068Z digest=sha256:fdf6c5a2570904faeae29936a464bd3cfda8677e6e20d30ca819af174c6867cf

Observation 9d52d7e1-7c04-4e59-b9d1-d849e87e2b27 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.205357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.205357Z digest=sha256:00ccb07fab4f7a7979c55dc357eb009ce0399f2730ddb22a1fe9bab01b1c5d5b

Observation d549b25a-4dac-4f19-9750-9debf0272fd8 · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.208807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.208807Z digest=sha256:f4545d61dd34fe716708d29e79e8f9fed7ae070d7ca9b489e6f68e25d8730242

Observation dc3de447-b1b9-46ac-9c95-d0bd99d06b01 · outbound

This paper cites Kwok, and Yu Zhang.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Kwok, and Yu Zhang

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.958067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.212605Z digest=sha256:95ca1e503b8c88a49e21d9488ffea4349cf808ebdad4a8af5e0ae092ca2a004c

Observation dd154da8-3727-42be-82e9-396a68385a9b · outbound

This paper cites Moderating Illicit Online Image Promo- tion for Unsafe User-Generated Content Games Using Large Vision-Language Models.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Moderating Illicit Online Image Promo- tion for Unsafe User-Generated Content Games Using Large Vision-Language Models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.947447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.215937Z digest=sha256:4c40422e693f58239909725bbd807ba3633aeb8011502acb13e461d455d9af2a

Observation fafedff1-86aa-4ec7-a2ca-2309d30b53ac · outbound

This paper cites LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.219871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.219871Z digest=sha256:9da44c14bbbc4e6f91149705a8673df7a0a70dde9a1fa3d2dcff503ed0adf25b

Observation fe8159af-8419-41c0-b7a7-1a5e90c57591 · outbound

This paper cites Glass, Akash Srivastava, and Pulkit Agrawal.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Glass, Akash Srivastava, and Pulkit Agrawal

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.936855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.224106Z digest=sha256:1ae17bca94f336e45f260b5fe1a32231ed81b37d791fc42d280b35ac9df674a3

Observation c9250c89-11de-428e-a8f1-c7800e71b84c · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.925456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.227446Z digest=sha256:24c217f0e7e99b5e99117e1a3e88312118c7c8ce2489a33cfd490adb2425d8d4

Observation a18f95d8-9b31-4586-a393-65111424bd3d · outbound

This paper cites Deep Reinforcement Learning for Di- alogue Generation.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Deep Reinforcement Learning for Di- alogue Generation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.915021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.231261Z digest=sha256:deccc859374d70eb49ec68f5436ef771a6da84a31998af6baeedc2696bbc4a3c

Observation 438ccd36-c875-413f-a23b-e4f14297bd45 · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.235072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.235072Z digest=sha256:bf6d60a09dae442aa50d928943fb9512df985e065348f62230623a0498709745

Observation 197ee628-b76b-4860-9504-86827db0aa01 · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Silkie: Preference Distillation for Large Visual Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.238949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.238949Z digest=sha256:031a8a1919ea51b1421d9d071f6c189e66dfab9ebcdbabb1a951f4e166b33300

Observation bea15e90-d2cf-4593-adcf-9b1e5939ef22 · outbound

This paper cites Red Teaming Visual Language Models.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Red Teaming Visual Language Models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.904973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.242675Z digest=sha256:2ae71bb8b1200aa158652872969bf2bb18b03d315b60a5bfad7af46dd64a14d7

Observation 9a707ee9-00d6-4970-a3c2-35751f12fba4 · outbound

This paper cites GOAT-Bench: Safety Insights to Large Multimodal Models through Meme-Based Social Abuse.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities GOAT-Bench: Safety Insights to Large Multimodal Models through Meme-Based Social Abuse

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.246231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.246231Z digest=sha256:0489ddc086080028f724dac3ba3c91483bb65d4674a6c3a4d5a1c0c0a8376714

Observation 3e7297ca-769d-49f1-9d06-003aa41fe1be · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.893067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.250252Z digest=sha256:ca688ba6e331a40e502c64a6d96705da732f627516cf10e7ea1f50ae8fad26d0

Observation b1f2be79-c6f2-4023-bd69-e7bf4bd24fca · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Improved Baselines with Visual Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.254512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.254512Z digest=sha256:ca1b74ba7105d2b453cb178f9777f59c88d81c17444214a7867aa863d3827e83

Observation fca06dda-adac-42f5-a31c-5adfe6756e1f · outbound

This paper cites Visual Instruction Tuning.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Visual Instruction Tuning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.880351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.258141Z digest=sha256:84c3daaf2e6a5e62f40574184ad88492efe14cf5272874d5b2be1b1528cd221f

Observation 2faa3e67-f4e6-46ee-85f4-9463ca6e7bcf · outbound

This paper cites Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.262536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.262536Z digest=sha256:f58359ae82cb818cf351d2c21fd5c6ee318f70cb9a7d16659863a758a64a49e4

Observation 5e894b5c-b5e4-4f10-99a8-b270c72bb9e9 · outbound

This paper cites MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.869284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.266120Z digest=sha256:68e2a586be4c1c80fd4b0deb51a95e1bda1b2d9dc09d811fe1749cd84bc6cbc4

Observation f03aac4f-842b-40ae-84fa-8b4b78a6425d · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.269179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.269179Z digest=sha256:16e9e5c9bc530203a1b1b03e228a2c30c89e056f4ccfc2f8825c62dddcec75b2

Observation 3acacffc-bf8e-4d23-8f7e-673730b424ca · outbound

This paper cites Safety Alignment for Vision Language Models.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Safety Alignment for Vision Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.272675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.272675Z digest=sha256:24c1d4934c7887ce01eef36f6464266147356d68e4369e6f018573bb0bb0af53

Observation a56f1111-94f5-472d-912c-ff112a1ca5a0 · outbound

This paper cites From Meme to Threat: On the Hateful Meme Understanding and Induced Hateful Content Generation in Open-Source Vision Language Mod- els.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities From Meme to Threat: On the Hateful Meme Understanding and Induced Hateful Content Generation in Open-Source Vision Language Mod- els

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.858852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.276398Z digest=sha256:ce79ac188eaa78252bdc9ecafb0c60a6128ed2bcd160e5ea66a0accdddb6ee0a

Observation 4e622c8d-4a8c-4177-b087-afdfaef88c04 · outbound

This paper cites an unresolved cited work.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:21:35.847374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.279490Z digest=sha256:f76138a7a672243e38cd89e7c3796e7215838f1a2f4168ca7712f0d0f6335459

Observation 50bc509b-1ddb-4157-850f-f19e99d996b9 · outbound

This paper cites GPT-4 Technical Report.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities GPT-4 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.282837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.282837Z digest=sha256:dfa045ca5c7bd2e26a7ce54c4f48dbd35e1d04067544b37d8a39e5f44fa7a0f9

Observation aa6b75b0-5afb-4771-8d0d-446cfaead8b1 · outbound

This paper cites Efros, and Trevor Darrell.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Efros, and Trevor Darrell

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.837911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.286650Z digest=sha256:b9d296025e8ee07ed16922bb34f0168068077e1dc22f917a8fa92ad3affeec4e

Observation b1632b72-81db-4f0f-861a-6e8f9ece354b · outbound

This paper cites Visual Adversarial Examples Jailbreak Aligned Large Language Models.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Visual Adversarial Examples Jailbreak Aligned Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.290281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.290281Z digest=sha256:b39e36b86708b16d877025530ab2a9fd323721204d2aa8adabf8fc9bca9bde79

Observation f67fa2e5-384f-4cc9-a0f4-9d271e4a1de4 · outbound

This paper cites On the Evolution of (Hateful) Memes by Means of Multimodal Contrastive Learn- ing.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities On the Evolution of (Hateful) Memes by Means of Multimodal Contrastive Learn- ing

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.828089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.293978Z digest=sha256:19fb3db5b3c72279013226534129d0164e76e8f77ebababeb8572dc9eb1401ee

Observation 011ea5f8-e16b-47fa-8bdd-fe77222e103f · outbound

This paper cites Unsafe Diffusion: On the Gen- eration of Unsafe Images and Hateful Memes From Text-To- Image Models.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Unsafe Diffusion: On the Gen- eration of Unsafe Images and Hateful Memes From Text-To- Image Models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.814657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.297189Z digest=sha256:98eb844dd0c4e4a2a83799181fb0e6fb5b1573d4c6acb867f8cb713001dfad5c

Observation 1aa2458a-4eb7-4ac5-8c6b-f652f18bcc17 · outbound

This paper cites UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.300886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.300886Z digest=sha256:22a559ef4820a7d8b9b6ff02df661f4cfadc1712f3b1a209f34581ab4e17510f

Observation da40a24f-6f79-47c9-8e59-5fcf662f07cf · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Learning Transferable Visual Models From Natural Language Supervision

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.802346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.304263Z digest=sha256:137e0f849a215a55ad0b312cd6e9b5777ae161321711e8be7ee2ef2ecff43f96

Observation 786f692e-4762-405e-8cda-7178cee67106 · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Manning, Stefano Ermon, and Chelsea Finn

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.789689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.307529Z digest=sha256:57d810e3dc12e55d6ddb3aaa679429f13e7e3853a6bb17d30f3211bac7eb1d6a

Observation 4c5e82af-a94f-420a-b7a3-d07d305c1978 · outbound

This paper cites Exploring the Limits of Zero Shot Vision Language Models for Hate Meme Detection: The Vulnerabilities and their Interpretations.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Exploring the Limits of Zero Shot Vision Language Models for Hate Meme Detection: The Vulnerabilities and their Interpretations

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.311066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.311066Z digest=sha256:d13f62b7cc8ace7f92e7b47a22e313eaec76cb672c0416bf0f600c57cb4b67dc

Observation c35be7c9-a6dd-4902-bcbc-86d09efac4a2 · outbound

This paper cites Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.315436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.315436Z digest=sha256:f6de3995849da6e17fef590219a5488382c25e29efb4e02b0173ff5184841886

Observation 3bb4527a-5233-4b6a-b959-237078f85288 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Proximal Policy Optimization Algorithms

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.319076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.319076Z digest=sha256:19a04c0c729f71a0285a35bf79a712a40c0f4feb86f9bf7577fc0e579edaab0c

Observation 197ef339-8d22-4290-808f-3a0a717836ca · outbound

This paper cites HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Cam- paigns.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Cam- paigns

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.322406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.322406Z digest=sha256:892727e0140c815689ab8527f7d7103f2c77b991a2859c42f0601484821f9c9d

Observation 3bfa66ed-6201-4f70-81cc-6ee2ef44e5e6 · outbound

This paper cites Assessment of Multimodal Large Language Models in Alignment with Human Values.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Assessment of Multimodal Large Language Models in Alignment with Human Values

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.326106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.326106Z digest=sha256:b37d3e4e6d8d44763b81e6d321a426aebbe8cf28e1956d6ddaf6653a308c5216

Observation f67551d3-b7b4-40ee-bf94-4af4e4236075 · outbound

This paper cites an unresolved cited work.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:21:35.768524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.329648Z digest=sha256:4fd5fe8f5549ae93a4f76e95552ab2b0d477ae883a5ce406150272ea3686c955

Observation fbd7e762-aab8-4f8d-a33f-cc4ad88a93c9 · outbound

This paper cites Align- ing Large Multimodal Models with Factually Augmented RLHF.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Align- ing Large Multimodal Models with Factually Augmented RLHF

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.757327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.333883Z digest=sha256:4910e307ca2a8f4f99e76ea07675b0694906b3eba42b678df0947791ebad52e4

Observation cfb9e0df-b385-4594-a38f-b907dc1c8ccb · outbound

This paper cites Vicuna: An Open-Source Chatbot Impress- ing GPT-4 with 90%* ChatGPT Quality.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Vicuna: An Open-Source Chatbot Impress- ing GPT-4 with 90%* ChatGPT Quality

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.745048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.337460Z digest=sha256:339e00f5f8e1b5ac0a547b601dd8c7c656a5e859d088fbe7d407b9fcd5d3a349

Observation d5218bbb-8a69-4f89-a888-3092290f110f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.340516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.340516Z digest=sha256:d3b27ac9845c648a4b96669e8fb2a1026143b913f5820993a66bdc39a47c923e

Observation 0e60edac-3f2c-4201-8a5d-126e1d465497 · outbound

This paper cites Safe Inputs but Unsafe Output: Benchmarking Cross-modality Safety Alignment of Large Vision-Language Model.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Safe Inputs but Unsafe Output: Benchmarking Cross-modality Safety Alignment of Large Vision-Language Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.344438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.344438Z digest=sha256:3d27758c299b06a46c2b5fb8fdc4102d92180b2f43c1ee3d2d49cde793263cdf

Observation d4c365a7-aa0e-4201-a726-517095b23028 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities CogVLM: Visual Expert for Pretrained Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.347771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.347771Z digest=sha256:6529404079b11532bae9ca78a125150c679d5a5cf5ed0dc5a8b43a87c35256f6

Observation b74250c5-fcb7-4e9e-bac2-d0389d5afa0d · outbound

This paper cites RL-VLM-F: Rein- forcement Learning from Vision Language Foundation Model Feedback.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities RL-VLM-F: Rein- forcement Learning from Vision Language Foundation Model Feedback

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.733568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.351460Z digest=sha256:95e0b586e02807b7e644d03d58785175d379cbe8f10c106fc653ed4ecaeda9d5

Observation b90fd3dc-1854-450a-b2cc-e1efdc81e7d4 · outbound

This paper cites The Perfect Blend: Redefining RLHF with Mixture of Judges.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities The Perfect Blend: Redefining RLHF with Mixture of Judges

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.355204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.355204Z digest=sha256:dcd10959ae95c7298609c6b96af939d0660b66b992027b3b3987973e2c5159f0

Observation 51af121b-4754-4101-9dfd-43e10568370e · outbound

This paper cites Qwen2 Technical Report.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Qwen2 Technical Report

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.359327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.359327Z digest=sha256:06410ee1dcbfe2c24585f5e606192ce1b5a866106a6753b2ea93c6e1c5c26feb

Observation be4f0882-b600-49f2-afe0-ca71eed6ddc9 · outbound

This paper cites Bridge the Modality and Capability Gaps in Vision-Language Model Selection.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Bridge the Modality and Capability Gaps in Vision-Language Model Selection

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:21:35.449506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.363566Z digest=sha256:4cfbbffae3d8dc80a74b9d5a672a98bad780b3f9bf1a22965e0cdf23f0adb91a

Observation 2118bebf-d9d8-4211-99e4-9b19d984e47d · outbound

This paper cites RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-Grained Correctional Human Feedback.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-Grained Correctional Human Feedback

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.721254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.367388Z digest=sha256:6f34680e8128dcefba6ba11c66f056805c93c12824fffc2715ee3dd91462f9f6

Observation ffc2c45d-0465-46fb-a516-7373399e78f1 · outbound

This paper cites SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.371849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.371849Z digest=sha256:ff43c09361a37ed4748245a64794f1a4f085dd864b726fa4d778a7c31ee0d2a8

Observation 4b83a515-e482-4667-93b7-ce63f42bb000 · outbound

This paper cites Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.375340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.375340Z digest=sha256:6f0c0d747d7de1798819b600a3e28acea6b573fb37b076cf3f310a3fd6f07399

Observation 641d9c1f-33ee-4821-a3c0-9a6db5d9a12c · outbound

This paper cites On Evaluating Ad- versarial Robustness of Large Vision-Language Models.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities On Evaluating Ad- versarial Robustness of Large Vision-Language Models

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.709325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.378803Z digest=sha256:2f4feb6a6cfe911aaa79eb2207d47cdf24350f3e452f3b40d5addc3c34a3f093

Observation 2c76d47f-024e-451e-a568-db48a2677b97 · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Secrets of RLHF in Large Language Models Part I: PPO

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.382133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.382133Z digest=sha256:c63e3f76e1d6625ffb8bd0842325df28a42b3432d6540b40d54079e4521b80df

Observation e19ae674-5361-406b-b663-9a91ebdf934b · outbound

This paper cites Yes” or “No.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Yes” or “No

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:35.697292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:21:35.386601Z digest=sha256:5c7d86d409030ec0b7bbb41f3ed8933df5be9259f2c8f8578721abf7f54320b6

Pith citing papers

No inbound Pith citation observations are available.