Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T14:10:16.760619Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:1908.04289.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T14:10:16.760619Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-14T12:22:25.530287Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-14T12:22:25.827597Z
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f8794b00-5993-4e9f-9142-22eebf63d381 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Bottom-up and top-down attention for image captioning and visual question answering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3941720b-7a5b-43cb-80ca-b9bd5042be37 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Vqa: Visual question answering
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 52832828-3328-471d-bdde-897415fba561 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Mutan: Multimodal tucker fusion for visual question answering
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d7fc90f7-b5cc-4559-9906-8d368a04006d · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63c24c3a-8e2f-48bf-b099-1c4c1ef00846 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Imagenet: A large-scale hierarchical image database
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 61c023c2-b3af-4047-95df-fd6490f84d96 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7df4cfc-ad11-4d94-9eea-30cb792f96f3 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 532a131c-4239-4c17-956d-5d8bde91264d · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Dy- namic fusion with intra-and inter-modality attention flow for visual question answering
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f626cac9-74e9-4d54-8c40-7c3694423d9b · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Question-guided hy- brid convolution for visual question answering
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8153b2b5-04b5-446e-a30f-0a7f0c59ab12 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Compact bilinear pooling
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 90176aa6-a73a-4583-9813-cbf714d5130b · outbound
Multi-modality Latent Interaction Network for Visual Question Answering 2nd place solution to the gqa challenge
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d74ce0eb-9090-44eb-9b36-e42a761c56a5 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 88e27c44-0f68-48ef-8eb1-1cef736b7665 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Deep residual learning for image recognition
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6a995f6b-c0bd-433d-a828-a4f9984090a5 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Relation networks for object detection
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4fcf0b6b-17f1-4c3c-acad-41557ff3f9be · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Weakly-supervised Compositional FeatureAggregation for Few-shot Recognition
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 478c626e-c65b-4415-a993-59441475868e · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Learning to segment every thing
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c800019c-d2df-4c7e-97fd-1769b900edfc · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Densely connected convolutional net- works
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ffb9b8b-e6e0-4ff7-aa07-2c926e53558b · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Video object detection with locally-weighted deformable neighbors
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e35de43b-3009-4694-9a00-0665efae6991 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Clevr: A diagnostic dataset for compositional language and elemen- tary visual reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7f53e586-a14c-4007-91de-fa7bf28b6dcc · outbound
Multi-modality Latent Interaction Network for Visual Question Answering An analysis of visual question answering algorithms
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c2d20e82-c705-4789-8e7b-4f4a32468a76 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Bilin- ear attention networks
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 789e25c9-c14a-47d9-9448-52a1758c08d1 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Hadamard Product for Low-rank Bilinear Pooling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7099b22-edbf-4701-875c-c95a3db858fa · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Adam: A Method for Stochastic Optimization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fba9b756-99a7-4e13-9aa4-b1a17e9f3783 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Skip-thought vectors
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6cfde9e3-b90f-4126-8fe4-a12cc0d4b3ed · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Imagenet classification with deep convolutional neural net- works
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 567d933e-59d2-4345-9993-f8dfe5c490fa · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Microsoft coco: Common objects in context
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b111dbf-8d54-49fd-9a3e-3b66434d39ef · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Improving referring expression grounding with cross-modal attention-guided erasing
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fe831b02-d01a-4554-b47e-7eee89b55662 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1e91f63-aeee-4e88-af30-73411c56a52c · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Hierarchical question-image co-attention for visual question answering
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 25361de9-a941-441c-b375-964800efefb6 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Distributed representations of words and phrases and their compositionality
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1a56c1bf-1028-4e01-a1ed-d84b7ac9c4cd · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cea23f11-e8c5-4b1f-a78f-7a5d3a403b19 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Training Recurrent Answering Units with Joint Loss Minimization for VQA
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60aacede-dd90-40e1-81e4-0b32d574e339 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Im- age question answering using convolutional neural network with dynamic parameter prediction
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 988bf229-828f-48e0-8335-df53c9ccadca · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Learning conditioned graph structures for interpretable vi- sual question answering
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9b0fb6fd-9f3c-4f04-9c4f-f11874844548 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Automatic differentiation in pytorch
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40a778a7-310d-451a-b131-44197d020028 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce22324b-860a-4eae-9220-d29dbf2e4eeb · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Glove: Global vectors for word representation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68e5d917-76b3-4481-877f-3c6e85b6a2b3 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Film: Visual reasoning with a general conditioning layer
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7522a3e6-cf18-4dde-80a8-d9cf8eb13b5c · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Deep contextualized word representations
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b48ce64b-0f92-4ee7-a237-1fec3d7f0918 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Language models are unsuper- vised multitask learners
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52362887-7d93-44d1-b0ee-0963074959e1 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Faster r-cnn: Towards real-time object detection with region proposal networks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 73f4e08b-e4aa-419a-9b1e-9953aa6767ac · outbound
Multi-modality Latent Interaction Network for Visual Question Answering A simple neural network module for relational rea- soning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2d6bca56-b516-466c-b929-64a53a53c5bf · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Question type guided attention in visual ques- tion answering
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 85f2cdf2-0d6e-4cbc-97fe-76967bd6ec6a · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Very Deep Convolutional Networks for Large-Scale Image Recognition
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6ee223f-13c4-4ed4-b049-003d9f6bf967 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Attention is all you need
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f9aa4ee0-2de4-4ff1-9ddb-b916780cc722 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Non-local neural networks
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e530c4c5-f0de-477b-8477-f737887aa4dd · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Pay Less Attention with Lightweight and Dynamic Convolutions
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1ff441c-aa0b-41ce-b19d-037395704ed9 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Show, attend and tell: Neural image caption gen- eration with visual attention
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 00a67c6a-a086-41c4-ad6c-cbc161570100 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Stacked attention networks for image question answering
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e82286dc-ef8a-4aab-aedb-93690a8ffbb2 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Scene Graph Reasoning with Prior Visual Relationship for Visual Question Answering
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7ec17cdb-760a-4643-b1c0-c87333ed883f · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Explor- ing visual relationship for image captioning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation feb6f77b-7d61-4ecb-b53b-76d295491612 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Beyond bilinear: generalized multimodal factorized high-order pooling for visual question answering
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation de7e3c54-6ffb-4a8d-a514-38cbd2ab6e8f · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Yin and Yang: Balancing and an- swering binary visual questions
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f4fcaa43-606e-477c-a945-f8fe8977609d · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Learning to Count Objects in Natural Images for Visual Question Answering
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57c5b259-ddc5-4643-99f5-a6f8b52f7753 · outbound
Multi-modality Latent Interaction Network for Visual Question Answering Structured attentions for visual question answering
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation abf6c8a0-e384-41c4-a8c4-1eeea82551cd · outbound
Multi-modality Latent Interaction Network for Visual Question Answering 2nd Place Solution to the GQA Challenge 2019
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 32e70082-8c07-4327-8b3d-e21b107212c7 · inbound
LXMERT: Learning Cross-Modality Encoder Representations from Transformers Multi-modality Latent Interaction Network for Visual Question Answering
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.