Pith. sign in

Paper Citation Record · LEDGER

BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 75 inbound Pith citation observations for arXiv:2201.12086.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2201.12086 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 75 of 75 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:27:13.254590Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

867
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation baa19c17-51eb-4465-be33-c1137286e3f4 · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.106705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:59e80e7de2c7af5d943a19e5dc8964d39273de1b799f59d73540574091368e1d

Observation bae1fd20-baed-4f46-b920-47083fb98875 · inbound

CoCa: Contrastive Captioners are Image-Text Foundation Models cites this paper.

CoCa: Contrastive Captioners are Image-Text Foundation Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:53:08.440780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T10:53:08.292063Z digest=sha256:970ae4857812ddf38512fe1e9ad5d757d260b78b509086ad4c1de7b4b2a5fed8

Observation 593de29b-f42b-43f9-a565-bfa484b3952f · inbound

LAION-5B: An open large-scale dataset for training next generation image-text models cites this paper.

LAION-5B: An open large-scale dataset for training next generation image-text models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:22:17.291001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T14:22:16.968028Z digest=sha256:c14ee92086ef72181241017a28469bbff8ec7b06ca563f837212e52ff0a4d416

Observation 9770d1f3-3e38-4f90-88bf-3591b6badb29 · inbound

Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models cites this paper.

Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T20:41:19.191871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:41:19.191871Z digest=sha256:258efd7dda2e418bfdbc3a82b7abf5a78cbd4f9a4c514abae19a365f90fd7ef8

Observation 2073b075-6195-4090-a649-687ad2ea368c · inbound

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning cites this paper.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.956799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.956799Z digest=sha256:0b40eea04d218930c6127903661edad760f9b51b354621767de79ca91a98c738

Observation 7f4ebab1-12e7-43f1-b7ab-17947b7f31d3 · inbound

VisGraphVar: A Benchmark Generator for Assessing Variability in Graph Analysis Using Large Vision-Language Models cites this paper.

VisGraphVar: A Benchmark Generator for Assessing Variability in Graph Analysis Using Large Vision-Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:52:57.048692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:52:57.048692Z digest=sha256:9802eab5b9d5507d629428d335697f699976627f9d7e91480a76a346ba28d571

Observation 89d19cbc-5e0b-4037-964f-f36f9e522909 · inbound

Health AI Developer Foundations cites this paper.

Health AI Developer Foundations BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T14:30:51.286288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:30:51.286288Z digest=sha256:c7295de49bc97da9191bed1173a979bd9a6dba611f21e84e55d31ee9f36b7a57

Observation 2e47c0ab-8447-4cf5-bb4f-4af55df8b0d1 · inbound

Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey cites this paper.

Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 241

Resolution
unresolved
no resolver link, observed 2026-08-12T12:02:31.527661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:02:31.527661Z digest=sha256:331793d28054c9351ca967b1c12791a548c323c852a3cfbde9cfea55ab24cbec

Observation eed1a4c1-c4c0-487e-837b-c911053a739d · inbound

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features cites this paper.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.674598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.674598Z digest=sha256:db41b7a42d206f3dfd03a082040d208ddb002cec648297b673625a9ea1e4f298

Observation 6847871d-1bde-49d8-a8ea-f8c81cab8a2c · inbound

SEMANTIC SEE-THROUGH GOGGLES: Wearing Linguistic Virtual Reality in (Artificial) Intelligence cites this paper.

SEMANTIC SEE-THROUGH GOGGLES: Wearing Linguistic Virtual Reality in (Artificial) Intelligence BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T23:16:56.050467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:16:56.050467Z digest=sha256:cb865f1a04c9b021b82baafd3c585fb2756cdf72d63ebee2709bbfaa17bcd5ce

Observation 99ed33be-fbd4-4c10-9f53-f5d89379b7c7 · inbound

SubstationAI: Multimodal Large Model-Based Approaches for Analyzing Substation Equipment Faults cites this paper.

SubstationAI: Multimodal Large Model-Based Approaches for Analyzing Substation Equipment Faults BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T05:51:35.234441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:51:35.234441Z digest=sha256:a391d1b64a48f47882579cdad0a4e1d945d79da60954c5e6c58a96b1cf30246c

Observation de06b389-afbe-44f8-826d-961e3a0560b2 · inbound

ZenSVI: An Open-Source Software for the Integrated Acquisition, Processing and Analysis of Street View Imagery Towards Scalable Urban Science cites this paper.

ZenSVI: An Open-Source Software for the Integrated Acquisition, Processing and Analysis of Street View Imagery Towards Scalable Urban Science BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T04:57:31.954129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:57:31.954129Z digest=sha256:1866386d75e833eeb351be06e44118566cfe75298ab5bad373f620f6bed73bea

Observation c198fae9-30b2-4257-ae32-808f2c83ef65 · inbound

ErgoChat: a Visual Query System for the Ergonomic Risk Assessment of Construction Workers cites this paper.

ErgoChat: a Visual Query System for the Ergonomic Risk Assessment of Construction Workers BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T23:51:00.774678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:51:00.774678Z digest=sha256:37e4c2596180c29172f4659047e65ecfc8d3aba3ef168dc6b05e76cb898e0bbc

Observation 4bfdaa4b-4bbd-4caf-9e20-00cb41d9d5b8 · inbound

Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment cites this paper.

Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:54.312338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:54.312338Z digest=sha256:8a93026ad89c47e8958c3278714e25806e43bb7b26d199c961c6129821c4dd8e

Observation bd1231c0-5bab-4062-a9a6-4704c070bda8 · inbound

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts cites this paper.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.508338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.508338Z digest=sha256:81fe2aea37051bf4f322e372ad225e56d020faca88093266b28a4587939051f7

Observation f4039f18-cc6a-4a87-b7a3-d2b57cfa79a7 · inbound

Visual Language Models as Operator Agents in the Space Domain cites this paper.

Visual Language Models as Operator Agents in the Space Domain BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:43.917693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:43.917693Z digest=sha256:562c1315891be86d9ca982bd302bd5947927fc9dedf6ee76c85077dffafca93e

Observation 5a6e88d5-6b6b-4b5e-8c8d-597058c38908 · inbound

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models cites this paper.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.159521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.159521Z digest=sha256:688c3730529ea63d5ae2d38a5d3d756e5089db9398ef6cb0e92b95561143110c

Observation d1958c10-42ea-438b-8eb3-9c9d01be1d5a · inbound

How Do Generative Models Draw a Software Engineer? A Case Study on Stable Diffusion Bias cites this paper.

How Do Generative Models Draw a Software Engineer? A Case Study on Stable Diffusion Bias BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:21.494606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:21.494606Z digest=sha256:9be9867639cb19c7978eb71e72e350977d0b3e48a0e410d3b5f0e18fd67ea260

Observation 0c9aec8b-a7cc-442d-bb24-7753da6d495a · inbound

Lossy Compression with Pretrained Diffusion Models cites this paper.

Lossy Compression with Pretrained Diffusion Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T19:44:31.705445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:44:31.705445Z digest=sha256:acb911568f0a663ead5984095f5c0db262f1e8abdf56c50dc20684217bfdbbe9

Observation 3f54cef2-46f8-4468-b683-f828dd1ac61d · inbound

StreamingRAG: Real-time Contextual Retrieval and Generation Framework cites this paper.

StreamingRAG: Real-time Contextual Retrieval and Generation Framework BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T15:24:49.345180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:24:49.345180Z digest=sha256:eda86d266c49ff187ef096830f05ef8a4d7cc97091f1dfea26875b1c8cb82e93

Observation 887e33d7-f74e-4daa-94d3-23655e2efce0 · inbound

Large Models in Dialogue for Active Perception and Anomaly Detection cites this paper.

Large Models in Dialogue for Active Perception and Anomaly Detection BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T13:36:47.024939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:36:47.024939Z digest=sha256:ada0f94c060455cc73d832b5ca87b676092456baca2fc5b403ba9a43c59fb7ac

Observation f10a70b6-0a7e-48e1-a9ce-3aae5d7e0072 · inbound

Generative AI for Vision: A Comprehensive Study of Frameworks and Applications cites this paper.

Generative AI for Vision: A Comprehensive Study of Frameworks and Applications BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T00:58:34.749290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:58:34.749290Z digest=sha256:4a4643d7d3eacd5bbdd0bba1ec3b248a4eaf030760d0bcb33f7bb58f52bdae77

Observation eca4d9dd-b8b6-4138-96d8-23b80e962e9d · inbound

RLS3: RL-Based Synthetic Sample Selection to Enhance Spatial Reasoning in Vision-Language Models for Indoor Autonomous Perception cites this paper.

RLS3: RL-Based Synthetic Sample Selection to Enhance Spatial Reasoning in Vision-Language Models for Indoor Autonomous Perception BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T22:09:40.627623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:09:40.627623Z digest=sha256:d27a522c48e0cedb4d35385796ca618b08edd43c88e735ce23e0b707df212f12

Observation 3706df61-7962-400b-b628-d8fafbf11f26 · inbound

Target-Augmented Shared Fusion-based Multimodal Sarcasm Explanation Generation cites this paper.

Target-Augmented Shared Fusion-based Multimodal Sarcasm Explanation Generation BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T12:58:24.178937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:58:24.178937Z digest=sha256:99709383310bfabe36a8766c1fafc009a8069855760b5267871a185daa54c521

Observation cac67ee3-c85f-4987-94ab-a04fdde58d7d · inbound

NanoVLMs: How small can we go and still make coherent Vision Language Models? cites this paper.

NanoVLMs: How small can we go and still make coherent Vision Language Models? BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.092936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.092936Z digest=sha256:774c6df7f3f65484dc961c242096c1b219d42b853cf67df9357dd2fc7aec9945

Observation 54af3713-1362-41c9-b479-7caf95bea41e · inbound

The AI-Therapist Duo: Exploring the Potential of Human-AI Collaboration in Personalized Art Therapy for PICS Intervention cites this paper.

The AI-Therapist Duo: Exploring the Potential of Human-AI Collaboration in Personalized Art Therapy for PICS Intervention BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T20:38:22.848810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:38:22.848810Z digest=sha256:6d6a14be4c972d60063c9cc790cdbd2f4b4ef980264f4778a2443eac781085f7

Observation 75f5a4ab-e5bf-4ba6-bc12-dcd437da3911 · inbound

Image Embedding Sampling Method for Diverse Captioning cites this paper.

Image Embedding Sampling Method for Diverse Captioning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T19:25:33.743835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:25:33.743835Z digest=sha256:f71ae7d2810cc88ac617c09c19f2bd52c1326e4492421f57ff1631e08c632519

Observation e08b8430-0906-43d4-b9e3-e5f9ed5b5033 · inbound

Optimizing Multi-Round Enhanced Training in Diffusion Models for Improved Preference Understanding cites this paper.

Optimizing Multi-Round Enhanced Training in Diffusion Models for Improved Preference Understanding BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T10:27:13.254590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:27:13.254590Z digest=sha256:82fe426ccdd275f4f76fef5260dc56858f4c6fbef2a7e4180aa17d166526a8d5

Observation 42146ad3-b134-4b69-961b-dbbfd3f5cb50 · inbound

MemeBLIP2: A novel lightweight multimodal system to detect harmful memes cites this paper.

MemeBLIP2: A novel lightweight multimodal system to detect harmful memes BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T05:13:23.796291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:13:23.796291Z digest=sha256:289940509feaa1e73e8b1b5ac2eccfd873680e3feba18d65ada9a58d5d7c9b2f

Observation 794f09dd-4141-4355-b892-a926bea30d99 · inbound

Multi-Modal Language Models as Text-to-Image Model Evaluators cites this paper.

Multi-Modal Language Models as Text-to-Image Model Evaluators BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:11.313155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:11.313155Z digest=sha256:176a5d76748e226b204b5ceefa5190323e8bfe097bb79efad6ac6efb5069fe63

Observation a3320923-bfec-449d-a15b-78ba96ab9fad · inbound

Mitigating Group-Level Fairness Disparities in Federated Visual Language Models cites this paper.

Mitigating Group-Level Fairness Disparities in Federated Visual Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:12:39.307006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:12:39.307006Z digest=sha256:1d185907fa81c04b18752ac99ff2c0756a3d0fcad55a94bb97ee3da23f16bca0

Observation 8d4e718e-278e-472a-948d-73f87d1af907 · inbound

A Vision-Language Model for Focal Liver Lesion Classification cites this paper.

A Vision-Language Model for Focal Liver Lesion Classification BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:39.551167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:39.551167Z digest=sha256:ec5702a9b9230487fa0e3f3fc6129659e54fb1bd06c31285b210d03582232373

Observation 2ec65496-e636-4164-a889-d5a25d926e84 · inbound

Multi-modal Synthetic Data Training and Model Collapse: Insights from VLMs and Diffusion Models cites this paper.

Multi-modal Synthetic Data Training and Model Collapse: Insights from VLMs and Diffusion Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.533992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.533992Z digest=sha256:ae7c4752706caec63d74d8f89a3383723b805f1bb16a41d2f3cd4e97740f4139

Observation f1a492cd-2259-4e83-9e96-aa9c3b0ec6b9 · inbound

GeoVLM: Improving Automated Vehicle Geolocalisation Using Vision-Language Matching cites this paper.

GeoVLM: Improving Automated Vehicle Geolocalisation Using Vision-Language Matching BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:14:56.317997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:14:56.317997Z digest=sha256:c88d2878c1b35a2e75f9b2437797dfbeb61b95d402165407f26b230e8586ec67

Observation b12929d5-5262-434c-84bd-2cdb3d5281e5 · inbound

Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method cites this paper.

Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:39.298701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:39.298701Z digest=sha256:3a54cae5119cd25adc211d530da404378724540e0b096abf651a32bffd4d3ff5

Observation 67f12da4-7886-4733-93d2-73a475b83fd1 · inbound

SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving cites this paper.

SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:27.296305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:27.296305Z digest=sha256:e9dd4328a9d6d722c4b57e307ae6375eb3608c4cb84ba2e74eb42dcb5131dacb

Observation 16464c40-9da1-4e3c-a1af-1ab136a3a7b7 · inbound

Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language Model cites this paper.

Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language Model BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:56.621484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:56.621484Z digest=sha256:f1aed8551ddf5c6bd0b8dbd5da01f0d31e5eeecda2beaac11fd7f42a09906ce5

Observation a14b41b1-eede-4c43-b21c-bf8d2717c0e5 · inbound

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction cites this paper.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.386112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.386112Z digest=sha256:63c800cec478883fe921fe94963a9935eda18cf85dc649d8914d93edf8e55a61

Observation 99c99c4c-5268-495e-8743-391df5f82fe2 · inbound

Seamless and Efficient Interactions within a Mixed-Dimensional Information Space cites this paper.

Seamless and Efficient Interactions within a Mixed-Dimensional Information Space BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 183

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:30.908336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:30.908336Z digest=sha256:06a09e22f17de1a45318b7c41a4be4bb3a622a4c764cb53591ddbb8f07313df4

Observation e3bc46e9-c2d9-4ca7-a005-b8171900903c · inbound

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge cites this paper.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:19.932135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:19.932135Z digest=sha256:7b579d7124a4bdb305aaaf098bcd8c8bbb0008307d21c89b9487c818a14d37af

Observation ea5a9fc6-5d0c-4dcb-8641-bf866810c160 · inbound

EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models cites this paper.

EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:22.467423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:22.467423Z digest=sha256:4280cce1779eae305dc224d14ff07938d747a6dcd6458d162ce9ec0369beeea1

Observation 42c0f4e7-0bd0-4110-a0fb-01d06fc962e2 · inbound

Graph-MLLM: Harnessing Multimodal Large Language Models for Multimodal Graph Learning cites this paper.

Graph-MLLM: Harnessing Multimodal Large Language Models for Multimodal Graph Learning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:14.835920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:14.835920Z digest=sha256:a85c439f08f0e3bb2ff650534fc635f5463d72a933f197397ca6cb089924bcb5

Observation db04e06f-f11b-45e7-a857-2b6f14e0b7e2 · inbound

CF-VLM:CounterFactual Vision-Language Fine-tuning cites this paper.

CF-VLM:CounterFactual Vision-Language Fine-tuning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.943940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.943940Z digest=sha256:55c4b0e214c6912db47500086e68a15c664832fd3d3d5ec9627e4d04dc4bd538

Observation 71e1d68b-15df-4b44-920d-feab10009b62 · inbound

AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation cites this paper.

AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:03:01.759523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:03:01.759523Z digest=sha256:02ced199c9793ebf3e425b3827c3d1fbb40ce576779e4dead66909f1f42d47d6

Observation 82005ae4-b23d-42ac-82a1-731b1d9e7fdf · inbound

CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning cites this paper.

CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:39.134498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:39.134498Z digest=sha256:40a67b6a5293cb6e322e24d25f55d8148a80e4616b7f0ed1913dbd0518a62164

Observation 8f1ccedf-c89b-40bd-a1e3-657f41d1ce3c · inbound

Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning cites this paper.

Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:52.700190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:52.700190Z digest=sha256:9cd64d257e05ded674bfddd14169d5beff07c430c626cace4c68af581107a081

Observation 0087c442-8ec5-4d59-879e-a0b2f32f4ba8 · inbound

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models cites this paper.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.813687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.813687Z digest=sha256:d977b42db133757a3ca0a24eb86307a0538c55a84eca8dd7981bb125f43089dc

Observation 438ccd36-c875-413f-a23b-e4f14297bd45 · inbound

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities cites this paper.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.235072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.235072Z digest=sha256:cdbc5800e0593411fc2a6065b3f8c673fa4b6ef001e7b1e8da48f2bdb676e1e4

Observation b1b5bbc1-8c36-47f8-a009-6567123d5122 · inbound

PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement cites this paper.

PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T16:22:49.656219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:22:49.656219Z digest=sha256:574c9f989e1ae94727174ace9f9b1a79b46220029f322f7b64de7be08280468f

Observation 5ae2999d-18ae-4e94-a659-52a64351a8ac · inbound

Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation cites this paper.

Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:08:35.573726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:08:35.573726Z digest=sha256:a82aae0aa2766a6af58aa10a09d86697926c2e3432b79c46f7d16806835a5d21

Observation 20cd51c6-8fcc-41d1-ab1d-36b60a47c7b5 · inbound

E.A.R.T.H.: Structuring Creative Evolution through Model Error in Generative AI cites this paper.

E.A.R.T.H.: Structuring Creative Evolution through Model Error in Generative AI BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:19.468148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:43:19.468148Z digest=sha256:664e966a0bd681c0658f34107edbbd6f61cda95addabfcfc4d79c2fea4cedefa

Observation 2afceee4-e63c-4777-8634-79791c7c625c · inbound

Affect-aware Cross-Domain Recommendation for Art Therapy via Music Preference Elicitation cites this paper.

Affect-aware Cross-Domain Recommendation for Art Therapy via Music Preference Elicitation BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T16:21:59.741127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:21:59.741127Z digest=sha256:511424c3ee3fabb3b11567fbe4b2164bd722d3aa6fa70f847bb430423e9401c4

Observation 546f64ed-bc4f-46f2-b00c-a17c7af023ba · inbound

Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding cites this paper.

Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T11:54:43.153675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:54:43.153675Z digest=sha256:52dbe2293b3a34df7d78f58f67fec155a1a64809c207a9738d29e46efd12e012

Observation ed8c2d38-64ab-4837-a083-f34616d960c3 · inbound

Visual Language Models as Zero-Shot Deepfake Detectors cites this paper.

Visual Language Models as Zero-Shot Deepfake Detectors BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:25.877668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:45:25.877668Z digest=sha256:cd0ead946d4a145a9dc4a692183050131c322d7709ff127831f3f787478d703e

Observation ab77093d-a728-44e0-b59a-2389c8319f8e · inbound

Investigating the Invertibility of Multimodal Latent Spaces: Limitations of Optimization-Based Methods cites this paper.

Investigating the Invertibility of Multimodal Latent Spaces: Limitations of Optimization-Based Methods BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T11:12:41.943831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:12:41.943831Z digest=sha256:c51e6f4b21599a6602b365903091aa28538e855f1833d7bb487ec21432da1b46

Observation 3056d7ab-380b-44a3-b4ea-be76eebf5597 · inbound

LEARN: A Story-Driven Layout-to-Image Generation Framework for STEM Instruction cites this paper.

LEARN: A Story-Driven Layout-to-Image Generation Framework for STEM Instruction BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:10:57.727313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:10:57.727313Z digest=sha256:1d3c19d5bc6b188b8604be005b3bb542df5b986ffc468e1331a55d2be8574351

Observation fb579a1d-9280-45a7-b825-c7649403e83b · inbound

UniECS: Unified Multimodal E-Commerce Search Framework with Gated Cross-modal Fusion cites this paper.

UniECS: Unified Multimodal E-Commerce Search Framework with Gated Cross-modal Fusion BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:15:20.724340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:15:20.724340Z digest=sha256:3ddb198b5fa1bf2084916f783e433baf9b2fdd3fa53ec3853cb2290f64d431b3

Observation 1d3e5fda-c22c-4f56-903e-f830935e267a · inbound

CLARIFY: A Specialist-Generalist Framework for Accurate and Lightweight Dermatological Visual Question Answering cites this paper.

CLARIFY: A Specialist-Generalist Framework for Accurate and Lightweight Dermatological Visual Question Answering BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T16:31:34.309538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:31:34.309538Z digest=sha256:560f62f7530e94a2fb72abedd5874742ae69002df689dc57c76c315169d8c320

Observation 655e029f-4504-4828-b44b-460328908220 · inbound

Effectively obtaining acoustic, visual and textual data from videos cites this paper.

Effectively obtaining acoustic, visual and textual data from videos BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.801004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.801004Z digest=sha256:e8c4673e005397d778a32990cb091fae903bf8938d8fc73d251ca58720602f83

Observation 41c23dbc-f4c8-4775-a078-abb7d87613f1 · inbound

Testing chatbots on the creation of encoders for audio conditioned image generation cites this paper.

Testing chatbots on the creation of encoders for audio conditioned image generation BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.951332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.951332Z digest=sha256:450a95bfec5dc02611a82a26a00c913c1aaf8b42499807163de1737477dda122

Observation 61c6fe04-d733-4125-b9a2-cb1e1d278238 · inbound

A Unified Geometric Space for Topological Alignment Between Transformer-Based Models and Human Brain Networks cites this paper.

A Unified Geometric Space for Topological Alignment Between Transformer-Based Models and Human Brain Networks BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T15:46:43.290470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:46:43.290470Z digest=sha256:8da1e4a2e79c90f3396c48be408543ec97fa4633f696571f3b3eff83c2fb0eaf

Observation 11c336d3-959d-44cc-9a77-7db3680058be · inbound

From Evidence to Verdict: An Agent-Based Forensic Framework for AI-Generated Image Detection cites this paper.

From Evidence to Verdict: An Agent-Based Forensic Framework for AI-Generated Image Detection BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:05:39.015725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T02:04:49.742067Z digest=sha256:2879f9af882a691968b18bf28a1d5d6f02267b7249b963ac4b45ea6f008079e4

Observation b2e4e563-2655-4094-8662-b19d6a1d9943 · inbound

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning cites this paper.

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T19:41:32.979486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:41:32.979486Z digest=sha256:d334d1ae3572e14bba6fb2716d1a49830613eb8531f352c0d279b44caf30ad66

Observation 82e03fb2-6006-4174-af81-e88bbc74baa0 · inbound

Integration of Object Detection and Small VLMs for Construction Safety Hazard Identification cites this paper.

Integration of Object Detection and Small VLMs for Construction Safety Hazard Identification BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:20:44.268586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T19:20:29.951227Z digest=sha256:ef3fc590286fc2d9778816568f05e4d694d007f25eb5dacf6d4d794a833fa0d5

Observation c3e09fb9-415b-4a57-a46d-647ff857880e · inbound

AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis cites this paper.

AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:25:45.504493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T19:22:08.306946Z digest=sha256:bc035c85e96c9673d1d9a08fe4489d30b2aa109adc6a5dc14b3abb0885eb6846

Observation 341e3caf-34c9-43a7-9d22-9bd5d07907b8 · inbound

Embedding Arithmetic: A Lightweight, Tuning-Free Framework for Post-hoc Bias Mitigation in Text-to-Image Models cites this paper.

Embedding Arithmetic: A Lightweight, Tuning-Free Framework for Post-hoc Bias Mitigation in Text-to-Image Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:20:53.927787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T05:20:47.706677Z digest=sha256:79c16dce694fe2540aeb70c3fd2870624fab40d02e202782170182cbbb613f41

Observation 891678bb-4bed-4fe6-b490-199c6925d33a · inbound

Multilingual Training and Evaluation Resources for Vision-Language Models cites this paper.

Multilingual Training and Evaluation Resources for Vision-Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:38:43.188960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T05:11:00.461167Z digest=sha256:df933a1b8913bc919ddf2f8fe5d7b98080c15e08e4f4cc15682ea42fff01feb1

Observation 6d6ce162-132d-45d7-be0c-2a148a666db9 · inbound

Multilingual Training and Evaluation Resources for Vision-Language Models cites this paper.

Multilingual Training and Evaluation Resources for Vision-Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T12:30:59.881443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:4f8d6784882aacf125bded79d4bc5422acb1359d5ce73c8c8978cb8c9dddaf6a

Observation cd14dc1d-a1ca-45a7-978d-62e056e5a4ad · inbound

Multimodal Cultural Heritage Knowledge Graph Extension with Language and Vision Models cites this paper.

Multimodal Cultural Heritage Knowledge Graph Extension with Language and Vision Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:15.851655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T12:08:49.134704Z digest=sha256:83e77b138b812ee9528fe9c107a6185f5ea7d28de57ed0562c5d36fb3ed3d87a

Observation fc76c320-8bfe-4b6b-810d-29a832473dcb · inbound

Your Embedding Model is SMARTer Than You Think cites this paper.

Your Embedding Model is SMARTer Than You Think BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:04:06.374816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T23:56:28.774601Z digest=sha256:79b0d17ba06e11d7c6fcbe9affd4a87797360f881cdc0f5dccdb40b6215fb6cd

Observation 3dd66f37-397d-40bd-92c5-55920bfd4cf1 · inbound

OmniCD: A Foundational Framework for Remote Sensing Image Change Detection Guided by Multimodal Semantics cites this paper.

OmniCD: A Foundational Framework for Remote Sensing Image Change Detection Guided by Multimodal Semantics BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:13:16.026784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T08:03:27.451288Z digest=sha256:46c5d6e5fc1ed590146297eb74862c12a198a45fead1d5064a38a78d3a33ff6d

Observation 827aa6c2-e4df-4fb1-9d36-aa21469e321e · inbound

The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model cites this paper.

The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T04:33:53.163084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:33:53.163084Z digest=sha256:e243f02d56e38062def97850e239a0236ae995079967ae63c20f8142281a0366

Observation acaad2c7-d8e6-4322-b728-6dcd62b80b01 · inbound

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation cites this paper.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T23:59:16.410456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:59:16.410456Z digest=sha256:e7d5312b08d445b7ae003dc4cfde6db1b31bb1cdc5272bb88bc0e5c5690eee16

Observation 0ad61f7e-cc04-463a-a242-3cbd1623810a · inbound

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation cites this paper.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.529655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.529655Z digest=sha256:27a9ce9ca88e06a601fa2e96861a7fcb6b9f1942250710a23e11b2e14b11ced4

Observation ff6b00fc-c954-4d5a-b5a3-6c3758cf700e · inbound

Diff-ID: Identity Consistent Facial Image Generation and Morphing via Diffusion Models cites this paper.

Diff-ID: Identity Consistent Facial Image Generation and Morphing via Diffusion Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T01:59:33.144686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:59:33.144686Z digest=sha256:7314833080bacc6f9c5196fab923390443f3afac8e7ee88ce00f134f3f19a717