Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T13:55:26.757084Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 3 inbound Pith citation observations for arXiv:1908.04107.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T13:55:26.757084Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-14T12:22:25.662003Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T03:39:30.840356Z
70 of 70 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e9783f87-858d-4eb4-812b-65d09ad8df0b · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Multimodal deep network embedding with integrated structure and attribute information,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 70e901c3-cc0d-40bc-a28c-5f92950e3cc7 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Discrim- inative coupled dictionary hashing for fast cross-media retrieval,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f6859418-36ed-429d-b6ab-d99ffa1366d7 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Shared predictive cross-modal deep quantization,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8b5b2dbe-a68f-456a-b955-3ba975161532 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Show, attend and tell: Neural image caption generation with visual attention
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3c17dd8c-eeeb-4498-98a4-fcdeaf39fe57 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions From deterministic to generative: multi-modal stochastic rnns for video cap- tioning,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 24dfb2cb-350e-48d1-8b1c-330066e7bc33 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Vqa: Visual question answering,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a712be8a-2deb-440c-8c4c-31ffb95e8ab3 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Ground- ing of textual phrases in images by reconstruction,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8391a958-710e-4be9-8f55-50c76d8493db · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Neural Machine Translation by Jointly Learning to Align and Translate
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9c99a69-6c26-4d21-9df4-79d3556906e4 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Recurrent models of visual attention,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 76a48858-2880-49ab-904d-fddcbfed84a8 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Draw: A recurrent neural network for image generation,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 025187f2-3715-4aa0-b7cb-2d689b66bb17 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Attention to scale: Scale-aware semantic image segmentation,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cc9ad11a-415f-4503-a2dd-4b836808f174 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Effective Approaches to Attention-based Neural Machine Translation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b80ff55-9a49-41b6-9d8f-8b089392bd47 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Deep biaffine attention for neural dependency parsing,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 83ed30e9-de93-4058-95d8-f2258e795eb8 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions A Neural Attention Model for Abstractive Sentence Summarization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ead648f-1b4b-4cdd-b058-1cc5db435c55 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Stacked attention net- works for image question answering,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 24d92376-f32e-44b1-9559-3cda4b74118f · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Multimodal compact bilinear pooling for visual question answering and visual grounding,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 93828af5-dd98-46b0-99cb-bde6d74cd5c6 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Hierarchical question-image co-attention for visual question answering,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 79de54b6-0d38-4847-8c0d-9db88de40266 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Multi-modal factorized bilinear pooling with co-attention learning for visual question answering,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2217e021-5f31-4498-b185-6e8a8e8ec933 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Bilinear attention networks,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b0fe8842-fcf6-4eea-ba58-794f04ddea9e · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Improved fusion of visual and language representations by dense symmetric co-attention for visual question an- swering,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8d40dc6e-7186-4ff5-bb03-b5adeb46d1d9 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Attention is all you need,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e4dea28e-54d3-44f4-9553-ddbda29ecfb3 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c50333a2-4c3f-4f51-8dc1-1ac7acbcc3b1 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Non-local neural net- works,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4d6962c0-7f30-4f83-80b5-ac0d6d0eff6b · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Relation networks for object detection,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c4fdcde9-6287-4302-8336-8c7a46ddfa72 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Making the v in vqa matter: Elevating the role of image understanding in visual question answering,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 51ff0f0b-5de1-4252-830d-3b16362c1344 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6e577d9d-4dcb-4137-b9d6-1b817db4de5c · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Referitgame: Referring to objects in photographs of natural scenes,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f30f20a7-9f63-40b1-b93a-9a8f24d590bd · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Generation and comprehension of unambiguous object descriptions,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6f797fb1-9b33-429f-a636-b0831eafd354 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Simple Baseline for Visual Question Answering
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d147a9b-af1f-4fb7-82c0-5b22177aadf3 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Hadamard Product for Low-rank Bilinear Pooling,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6d9aaaf4-c355-4c10-ab2c-6d3ebe5326e1 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Mutan: Multi- modal tucker fusion for visual question answering,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6d91550f-15d5-4026-88b0-8a0b3148f76a · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions ABC-CNN: An Attention Based Convolutional Neural Network for Visual Question Answering
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 109b5f20-38ee-426c-8fee-c9af7d2abde6 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions A Focused Dynamic Attention Model for Visual Question Answering
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb9f24d8-e8e2-4819-8e40-6397ae70a6fa · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Where to look: Focus regions for visual question answering,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fb59863f-722b-4c6c-be73-b56a7b16b79a · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Beyond bilinear: Generalized multi-modal factorized high-order pooling for visual question answer- ing,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 960371a1-dd9e-41eb-aca8-324d0c3487c9 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions A joint speaker-listener- reinforcer model for referring expressions,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fd8e2279-07a3-4bb7-a782-0b2172e20416 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Edge boxes: Locating object proposals from edges,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation eaaee506-d5f9-455d-a353-4f846e9de4db · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Deep residual learning for image recognition,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 848abf49-fe14-48b4-8ed8-c3d6ab67e6a9 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Rethinking diversified and discriminative proposal generation for visual grounding,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 09816d98-c73d-422d-96a0-3633ac698948 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Mattnet: Modular attention network for referring expression comprehension,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 23a1ab66-d516-47ae-a85e-05211370ff20 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Parallel attention: A unified framework for visual object discovery through dialogs and queries,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9aee77a4-0b7a-4c8d-8d49-bcb2a792117e · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Visual grounding via accumulated attention,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 619e07f7-7290-4b83-a249-6b0e2840a833 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Be- yond rnns: Positional self-attention with co-attention for video question answering,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b487a113-c177-42c6-920b-73ddb6fbff3b · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8e1ba17d-fc08-41e7-b6a5-609798755313 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Improving language understanding by generative pre-training,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 56e1d26b-e8d1-4ce0-84c8-e33d878dfbfc · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Factorized bilinear models for image recognition,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7c91eab1-0e15-4cd8-a276-5bb7285988bc · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Layer Normalization
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c91039a-5194-42a8-a217-dc1fad636ef9 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Glove: Global vectors for word representation
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ca5f6cfb-f77c-4d5d-aa6c-22c2b451f2d7 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Long short-term memory,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0c6bc30-b4b8-4e0b-a17d-62af13d28368 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Bottom-up and top-down attention for image captioning and visual question answering,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a01fe0f2-a863-4ebd-9edd-2c4db59909f0 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Tips and Tricks for Visual Question Answering: Learnings from the 2017 Challenge
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f73072f-c372-4c77-9609-1f8bc3152ae5 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Fast r-cnn,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 59e2ed05-d7a5-4aba-9e46-0538c12b7244 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Microsoft coco: Common objects in context,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5617209c-ca98-4c12-938a-ea8847087550 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Adam: A Method for Stochastic Optimization
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccee27c2-86b0-4da9-b579-2767d2a9f8b1 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Faster r-cnn: Towards real- time object detection with region proposal networks,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 75b9dd9d-9b4e-436b-a1c0-6d5dae966155 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5d1f2f3-65f2-427c-a4ef-e1246434acb0 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Compositional Attention Networks for Machine Reasoning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b37c205-f5f0-4f0b-87de-189dacaa1a74 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Mask r-cnn,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dc40e05e-a0b8-42f8-8ecb-56e048881c88 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Modeling context in referring expressions,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1cfe4aa3-3e50-4d06-89ae-46bce9abedf8 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Learning to count objects in natural images for visual question answering,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5513a91a-6151-457e-b88c-d5b26a4e0e03 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Deep modular co- attention networks for visual question answering,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f5eae359-dab9-4bab-b5de-cd6fe415a587 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Learning to reason: End-to-end module networks for visual question answering,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a20f6b27-33c3-4b36-a99a-bdcd5ada6b5e · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions A simple neural network module for relational reasoning,
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 263fe6b5-0431-4cd3-9e01-869ad04fabc7 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Inferring and executing programs for visual reasoning,
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3982b222-2158-42f5-9eaa-2b9aa7fdce41 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Film: Visual reasoning with a general conditioning layer,
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fe3807c5-9063-4fe5-b5ee-88a6eb62baa5 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Ssd: Single shot multibox detector,
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cb4a4a63-7375-47e7-9e60-4868dba22dee · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Very Deep Convolutional Networks for Large-Scale Image Recognition
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 113d83f9-c6f5-44b9-9627-4a54f7bf7bbf · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Referring expression generation and comprehension via attributes,
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 056540f1-7b59-45d8-bbe2-4ec13864cf9e · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Modeling relationships in referential expressions with compositional modular networks,
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ce81e276-226a-4c4f-9822-e6f343ce6cf9 · outbound
Multimodal Unified Attention Networks for Vision-and-Language Interactions Grounding referring expressions in images by variational context,
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e801b09e-8839-475b-bff5-3d70b5134379 · inbound
LXMERT: Learning Cross-Modality Encoder Representations from Transformers Multimodal Unified Attention Networks for Vision-and-Language Interactions
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 727d9b13-5265-4449-ad2c-613515bb85a3 · inbound
ViASNet: A Video Ad Saliency Network for Predicting Dynamic Saliency and Viewer Engagement Multimodal Unified Attention Networks for Vision-and-Language Interactions
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 419bf000-2dcf-42f0-a8e1-6427010faba4 · inbound
Alzheimer's Disease Diagnosis using a Multimodal Approach with 3D MRI and PET Multimodal Unified Attention Networks for Vision-and-Language Interactions
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.