Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 75 inbound Pith citation observations for arXiv:2201.12086.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:27:13.254590Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
867
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation baa19c17-51eb-4465-be33-c1137286e3f4 · inbound
Flamingo: a Visual Language Model for Few-Shot Learning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bae1fd20-baed-4f46-b920-47083fb98875 · inbound
CoCa: Contrastive Captioners are Image-Text Foundation Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 593de29b-f42b-43f9-a565-bfa484b3952f · inbound
LAION-5B: An open large-scale dataset for training next generation image-text models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9770d1f3-3e38-4f90-88bf-3591b6badb29 · inbound
Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2073b075-6195-4090-a649-687ad2ea368c · inbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f4ebab1-12e7-43f1-b7ab-17947b7f31d3 · inbound
VisGraphVar: A Benchmark Generator for Assessing Variability in Graph Analysis Using Large Vision-Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89d19cbc-5e0b-4037-964f-f36f9e522909 · inbound
Health AI Developer Foundations BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e47c0ab-8447-4cf5-bb4f-4af55df8b0d1 · inbound
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 241
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eed1a4c1-c4c0-487e-837b-c911053a739d · inbound
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6847871d-1bde-49d8-a8ea-f8c81cab8a2c · inbound
SEMANTIC SEE-THROUGH GOGGLES: Wearing Linguistic Virtual Reality in (Artificial) Intelligence BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99ed33be-fbd4-4c10-9f53-f5d89379b7c7 · inbound
SubstationAI: Multimodal Large Model-Based Approaches for Analyzing Substation Equipment Faults BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de06b389-afbe-44f8-826d-961e3a0560b2 · inbound
ZenSVI: An Open-Source Software for the Integrated Acquisition, Processing and Analysis of Street View Imagery Towards Scalable Urban Science BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c198fae9-30b2-4257-ae32-808f2c83ef65 · inbound
ErgoChat: a Visual Query System for the Ergonomic Risk Assessment of Construction Workers BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bfdaa4b-4bbd-4caf-9e20-00cb41d9d5b8 · inbound
Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd1231c0-5bab-4062-a9a6-4704c070bda8 · inbound
Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4039f18-cc6a-4a87-b7a3-d2b57cfa79a7 · inbound
Visual Language Models as Operator Agents in the Space Domain BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a6e88d5-6b6b-4b5e-8c8d-597058c38908 · inbound
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1958c10-42ea-438b-8eb3-9c9d01be1d5a · inbound
How Do Generative Models Draw a Software Engineer? A Case Study on Stable Diffusion Bias BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c9aec8b-a7cc-442d-bb24-7753da6d495a · inbound
Lossy Compression with Pretrained Diffusion Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f54cef2-46f8-4468-b683-f828dd1ac61d · inbound
StreamingRAG: Real-time Contextual Retrieval and Generation Framework BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 887e33d7-f74e-4daa-94d3-23655e2efce0 · inbound
Large Models in Dialogue for Active Perception and Anomaly Detection BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f10a70b6-0a7e-48e1-a9ce-3aae5d7e0072 · inbound
Generative AI for Vision: A Comprehensive Study of Frameworks and Applications BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eca4d9dd-b8b6-4138-96d8-23b80e962e9d · inbound
RLS3: RL-Based Synthetic Sample Selection to Enhance Spatial Reasoning in Vision-Language Models for Indoor Autonomous Perception BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3706df61-7962-400b-b628-d8fafbf11f26 · inbound
Target-Augmented Shared Fusion-based Multimodal Sarcasm Explanation Generation BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cac67ee3-c85f-4987-94ab-a04fdde58d7d · inbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54af3713-1362-41c9-b479-7caf95bea41e · inbound
The AI-Therapist Duo: Exploring the Potential of Human-AI Collaboration in Personalized Art Therapy for PICS Intervention BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75f5a4ab-e5bf-4ba6-bc12-dcd437da3911 · inbound
Image Embedding Sampling Method for Diverse Captioning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e08b8430-0906-43d4-b9e3-e5f9ed5b5033 · inbound
Optimizing Multi-Round Enhanced Training in Diffusion Models for Improved Preference Understanding BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42146ad3-b134-4b69-961b-dbbfd3f5cb50 · inbound
MemeBLIP2: A novel lightweight multimodal system to detect harmful memes BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 794f09dd-4141-4355-b892-a926bea30d99 · inbound
Multi-Modal Language Models as Text-to-Image Model Evaluators BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3320923-bfec-449d-a15b-78ba96ab9fad · inbound
Mitigating Group-Level Fairness Disparities in Federated Visual Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d4e718e-278e-472a-948d-73f87d1af907 · inbound
A Vision-Language Model for Focal Liver Lesion Classification BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ec65496-e636-4164-a889-d5a25d926e84 · inbound
Multi-modal Synthetic Data Training and Model Collapse: Insights from VLMs and Diffusion Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1a492cd-2259-4e83-9e96-aa9c3b0ec6b9 · inbound
GeoVLM: Improving Automated Vehicle Geolocalisation Using Vision-Language Matching BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b12929d5-5262-434c-84bd-2cdb3d5281e5 · inbound
Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67f12da4-7886-4733-93d2-73a475b83fd1 · inbound
SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16464c40-9da1-4e3c-a1af-1ab136a3a7b7 · inbound
Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language Model BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a14b41b1-eede-4c43-b21c-bf8d2717c0e5 · inbound
RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99c99c4c-5268-495e-8743-391df5f82fe2 · inbound
Seamless and Efficient Interactions within a Mixed-Dimensional Information Space BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 183
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3bc46e9-c2d9-4ca7-a005-b8171900903c · inbound
From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea5a9fc6-5d0c-4dcb-8641-bf866810c160 · inbound
EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42c0f4e7-0bd0-4110-a0fb-01d06fc962e2 · inbound
Graph-MLLM: Harnessing Multimodal Large Language Models for Multimodal Graph Learning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db04e06f-f11b-45e7-a857-2b6f14e0b7e2 · inbound
CF-VLM:CounterFactual Vision-Language Fine-tuning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71e1d68b-15df-4b44-920d-feab10009b62 · inbound
AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82005ae4-b23d-42ac-82a1-731b1d9e7fdf · inbound
CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f1ccedf-c89b-40bd-a1e3-657f41d1ce3c · inbound
Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0087c442-8ec5-4d59-879e-a0b2f32f4ba8 · inbound
Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 438ccd36-c875-413f-a23b-e4f14297bd45 · inbound
Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1b5bbc1-8c36-47f8-a009-6567123d5122 · inbound
PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ae2999d-18ae-4e94-a659-52a64351a8ac · inbound
Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20cd51c6-8fcc-41d1-ab1d-36b60a47c7b5 · inbound
E.A.R.T.H.: Structuring Creative Evolution through Model Error in Generative AI BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2afceee4-e63c-4777-8634-79791c7c625c · inbound
Affect-aware Cross-Domain Recommendation for Art Therapy via Music Preference Elicitation BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 546f64ed-bc4f-46f2-b00c-a17c7af023ba · inbound
Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed8c2d38-64ab-4837-a083-f34616d960c3 · inbound
Visual Language Models as Zero-Shot Deepfake Detectors BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab77093d-a728-44e0-b59a-2389c8319f8e · inbound
Investigating the Invertibility of Multimodal Latent Spaces: Limitations of Optimization-Based Methods BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3056d7ab-380b-44a3-b4ea-be76eebf5597 · inbound
LEARN: A Story-Driven Layout-to-Image Generation Framework for STEM Instruction BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb579a1d-9280-45a7-b825-c7649403e83b · inbound
UniECS: Unified Multimodal E-Commerce Search Framework with Gated Cross-modal Fusion BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d3e5fda-c22c-4f56-903e-f830935e267a · inbound
CLARIFY: A Specialist-Generalist Framework for Accurate and Lightweight Dermatological Visual Question Answering BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 655e029f-4504-4828-b44b-460328908220 · inbound
Effectively obtaining acoustic, visual and textual data from videos BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41c23dbc-f4c8-4775-a078-abb7d87613f1 · inbound
Testing chatbots on the creation of encoders for audio conditioned image generation BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61c6fe04-d733-4125-b9a2-cb1e1d278238 · inbound
A Unified Geometric Space for Topological Alignment Between Transformer-Based Models and Human Brain Networks BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11c336d3-959d-44cc-9a77-7db3680058be · inbound
From Evidence to Verdict: An Agent-Based Forensic Framework for AI-Generated Image Detection BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b2e4e563-2655-4094-8662-b19d6a1d9943 · inbound
Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82e03fb2-6006-4174-af81-e88bbc74baa0 · inbound
Integration of Object Detection and Small VLMs for Construction Safety Hazard Identification BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c3e09fb9-415b-4a57-a46d-647ff857880e · inbound
AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 341e3caf-34c9-43a7-9d22-9bd5d07907b8 · inbound
Embedding Arithmetic: A Lightweight, Tuning-Free Framework for Post-hoc Bias Mitigation in Text-to-Image Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 891678bb-4bed-4fe6-b490-199c6925d33a · inbound
Multilingual Training and Evaluation Resources for Vision-Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6d6ce162-132d-45d7-be0c-2a148a666db9 · inbound
Multilingual Training and Evaluation Resources for Vision-Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cd14dc1d-a1ca-45a7-978d-62e056e5a4ad · inbound
Multimodal Cultural Heritage Knowledge Graph Extension with Language and Vision Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fc76c320-8bfe-4b6b-810d-29a832473dcb · inbound
Your Embedding Model is SMARTer Than You Think BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3dd66f37-397d-40bd-92c5-55920bfd4cf1 · inbound
OmniCD: A Foundational Framework for Remote Sensing Image Change Detection Guided by Multimodal Semantics BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 827aa6c2-e4df-4fb1-9d36-aa21469e321e · inbound
The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acaad2c7-d8e6-4322-b728-6dcd62b80b01 · inbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ad61f7e-cc04-463a-a242-3cbd1623810a · inbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff6b00fc-c954-4d5a-b5a3-6c3758cf700e · inbound
Diff-ID: Identity Consistent Facial Image Generation and Morphing via Diffusion Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.