Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T14:54:41.910373Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 29 inbound Pith citation observations for arXiv:1908.02265.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T14:54:41.910373Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:43:42.788616Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
45 of 45 outbound references displayed
External citation measurements
1675
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation b8fae9c0-6b9c-4390-b0be-a5f30ffae3f0 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1fa016a7-422d-4cba-8bd5-0e8e5ff667e9 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 405e9517-e5a1-4f83-a976-2fb3e3f4af00 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Lawrence Zitnick, and Devi Parikh
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7a95221d-8e15-4709-b46a-7c8810a23519 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3c67ad70-a369-49f9-99d6-58153b0cd641 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e75e8e42-e1ab-4f93-aa03-248e64581f24 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks foil it! find one mismatch between image and language caption
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 27ab7fd0-c77c-4ec6-93ce-0865b1f768ca · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Embodied Question Answering
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3589cdaf-866d-4745-9e10-7d0d480c7bb9 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d6e2e915-46ed-455d-947e-f9c7b94ce0ff · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Don’t just assume; look and answer: Overcoming priors for visual question answering
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d98e4f41-c2bc-4011-8377-39a3080dbbd4 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks nocaps: novel object captioning at scale
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6c76b031-dcbe-45a3-a7f7-899775b61797 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Deep residual learning for image recognition
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 729e08fb-c372-4219-8db9-5453e054af65 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bede3cb3-c9f0-4199-8aea-471431a0a59e · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c4977e9b-8f1f-483f-8838-3a155826a247 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Improving language understanding with unsupervised learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fd8d8339-eb10-4cec-a6a3-60aa5660ff73 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Berg, and Li Fei-Fei
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44c8df1b-c7ff-47d1-ae94-4e94ac6e09fd · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a931e69-6555-4cb6-bef2-a8d7bb5b2719 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cc52f0f3-3bb0-4761-8ec2-e57fad484d82 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks URL https://en.wikipedia.org/
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3d237df6-a9d9-4524-b513-dd37c38d4592 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks One billion word benchmark for measuring progress in statistical language modeling
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ea437fef-cd26-4a3b-8feb-dfc3f86daa18 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Colorization as a proxy task for visual understanding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 18dcce10-bc79-4977-9c3d-3252c8e6af42 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Shapecodes: self-supervised feature learning by lifting views to viewgrids
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dd8d5424-451d-49a5-b858-18ead5c46d8a · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Look, listen and learn
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7f047bb8-390d-4482-904e-be35fe57c304 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Learning features by watching objects move
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 99e38133-ad65-4215-9531-35d2080f0e34 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cebbe19-81d2-4467-a158-ebb002609938 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks From recognition to cognition: Visual commonsense reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd8edebe-f4bb-49fe-a9e0-129127aa1104 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aecd10d1-5847-46e2-9aa3-ba19f2aa93cd · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Attention is all you need
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac8f1a2f-9a10-47d5-bc4a-d0afdfe82cf4 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e96ad61-628f-4e10-8e03-2e6cf31489c0 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks VideoBERT: A Joint Model for Video and Language Representation Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14aab524-892e-41c8-affd-f3078120b476 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Bottom-up and top-down attention for image captioning and visual question answering
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fb83a2d-8745-4d42-9ed8-88462481760e · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Faster r-cnn: Towards real-time object detection with region proposal networks
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e857b7d4-7934-4897-8e8c-1273306c134b · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Referitgame: Referring to objects in photographs of natural scenes
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7dad33ac-6be7-455a-9c84-8837ddbc74ae · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Mattnet: Modular attention network for referring expression comprehension
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6a66df2f-075a-4e3c-b7ba-c46c540614d0 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Mask r-cnn
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7e59628b-d5b8-49b3-8cee-a0b6dfe14169 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Stacked cross attention for image-text matching
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c77166b2-c5d1-4e60-9b1d-7334467c6ed2 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 411ea18a-47e1-432f-b2df-a10478bbd2db · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1486f96-6943-4e04-b33e-8eddb4b02df5 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Unsupervised visual representation learning by context prediction
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 71a9d6fa-5b50-4038-8fb8-8bb3039410a6 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Colorful image colorization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e46e166b-584a-4b27-81d6-e61180a2d4b8 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Discriminative unsupervised feature learning with exemplar convolutional neural networks
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6be4d9c9-cc7c-43c8-8403-8b064124c98c · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Context encoders: Feature learning by inpainting
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e43164d-2b37-4ba6-9ef9-e83ad7cfeaa0 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Learning image representations tied to ego-motion
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 90800b48-2a7b-473a-ae9a-2bc9413ba1a5 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Shuffle and learn: unsupervised learning using temporal order verification
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1418ecdb-7534-4fc7-af88-9bd735f8adba · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Cross-lingual Language Model Pretraining
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfcc0816-f708-4cf1-88f2-24e2ba26c6a1 · outbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Courville
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ff96a687-5a8e-4fbb-971e-7ba9a2fcfec8 · inbound
VisualBERT: A Simple and Performant Baseline for Vision and Language ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fe831b02-d01a-4554-b47e-7eee89b55662 · inbound
Multi-modality Latent Interaction Network for Visual Question Answering ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22ca41b3-de21-4c32-925f-6933a0f7f51a · inbound
Fusion of Detected Objects in Text for Visual Question Answering ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ee6b0e0-0cef-45be-b6ad-9436b121f399 · inbound
Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18ccb43a-3cd4-4824-9e2b-eec14122377e · inbound
LXMERT: Learning Cross-Modality Encoder Representations from Transformers ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea78c704-dae8-4dad-b4a5-705cab9935d2 · inbound
VL-BERT: Pre-training of Generic Visual-Linguistic Representations ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dde256fc-9f6d-475a-9589-44e54120cb74 · inbound
Text and Code Embeddings by Contrastive Pre-Training ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2afcd137-29e2-48ca-ab18-d7b7ebbb7240 · inbound
A Comprehensive Survey on Visual Question Answering Datasets and Algorithms ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33a53c64-c86b-403a-a2be-a5762bb1a7d3 · inbound
VisGraphVar: A Benchmark Generator for Assessing Variability in Graph Analysis Using Large Vision-Language Models ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9d4fabe-2f81-400e-8c38-39e7c695f6e8 · inbound
SentiXRL: An advanced large language Model Framework for Multilingual Fine-Grained Emotion Classification in Complex Text Environment ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c76cf31f-1383-49d6-8c21-87bd74ed90a1 · inbound
Multimodal Multihop Source Retrieval for Web Question Answering ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d91dae2a-3456-4a2f-ad03-b7eae883678c · inbound
Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 883a7e11-56e0-4ad0-bde5-c29ede8dd36a · inbound
The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b875199-5513-4653-99ad-c25625dfa6a5 · inbound
Performance Analysis of Traditional VQA Models Under Limited Computational Resources ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87ea1a9c-bf48-4b70-a47a-df80579b30b7 · inbound
Vision-Language Models for Edge Networks: A Comprehensive Survey ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f6bb367-4920-4ce2-9860-cea82207a6a0 · inbound
Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16d433d1-594f-41b7-90f4-22b6da9fb43c · inbound
Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18d9656d-e6ea-45f0-9bac-d593c441cf48 · inbound
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8519c36c-eef2-429e-aa49-9d5a12ae840e · inbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18a56846-eb1f-4f9e-8a8b-53be205dc05b · inbound
On the Resilience of Underwater Semantic Wireless Communications ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 233de5ba-3820-4e13-b61c-6c4c11479083 · inbound
Representations in vision and language converge in a shared, multidimensional space of perceived similarities ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 6241
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 691bee43-3208-4384-888e-71ed7341d524 · inbound
AME: Aligned Manifold Entropy for Robust Vision-Language Distillation ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e585f74d-cd12-4a2c-a89d-fea58f28007f · inbound
From Image Captioning to Visual Storytelling ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96d59a51-e9ae-4059-99ad-9723eee7f085 · inbound
Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 09885d11-81f9-41f3-bb38-1cec11183d43 · inbound
Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4dfd39e3-fc51-48b1-9eb4-04611f821c0f · inbound
Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 878188b0-7386-4b73-9b97-e531469b0240 · inbound
KG-TRACE: A Neuro-Symbolic Framework for Mechanistic Grounding in Antimicrobial Resistance Prediction ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e1deed52-e357-4079-9b7e-30134d9c6f08 · inbound
A Comparison of Fusion Techniques for Multi-Modal Human Activity Recognition on the HARMES Dataset ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c2c49c0a-a86a-4b2d-82cd-01515858c4fb · inbound
When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 112
Source-reported events for the cited work
Unavailable: canonical work link unavailable.