Pith. sign in

Paper Citation Record · LEDGER

Multi-modality Latent Interaction Network for Visual Question Answering

As of 15 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:1908.04289.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.04289 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T14:10:16.760619Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:22:25.530287Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-14T12:22:25.827597Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact3
  • verified fuzzy35
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f8794b00-5993-4e9f-9142-22eebf63d381 · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Bottom-up and top-down attention for image captioning and visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.556713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.516954Z digest=sha256:0dafa04fc9c8a924ba8159ad358cbfbbd296acc2cd738cb79111ff6929c413f3

Observation 3941720b-7a5b-43cb-80ca-b9bd5042be37 · outbound

This paper cites Vqa: Visual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Vqa: Visual question answering

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.542685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.522005Z digest=sha256:6625b4d2576905fba1564fe1aa34633973a23f3d8498bd227918314f8d6734a4

Observation 52832828-3328-471d-bdde-897415fba561 · outbound

This paper cites Mutan: Multimodal tucker fusion for visual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Mutan: Multimodal tucker fusion for visual question answering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.527839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.526459Z digest=sha256:7167e545a79039cd4245b9a1bbe3138eb994c235e42c03405125242f03c002fe

Observation d7fc90f7-b5cc-4559-9906-8d368a04006d · outbound

This paper cites Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning.

Multi-modality Latent Interaction Network for Visual Question Answering Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.531280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.531280Z digest=sha256:6e301afc1ce76bc610132aed831ac49f49f049cbf15a4b449143749767472067

Observation 63c24c3a-8e2f-48bf-b099-1c4c1ef00846 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Multi-modality Latent Interaction Network for Visual Question Answering Imagenet: A large-scale hierarchical image database

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.506296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.535373Z digest=sha256:879fd8e2495f48c558948f8c5b89fbc780e67c0f213d7944880e1c8009fc69c2

Observation 61c023c2-b3af-4047-95df-fd6490f84d96 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Multi-modality Latent Interaction Network for Visual Question Answering BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.539555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.539555Z digest=sha256:d49b1bbaa3a5b509c03ed2d1318a7edd37170be54da84e5c3fa0ba913a51a55f

Observation c7df4cfc-ad11-4d94-9eea-30cb792f96f3 · outbound

This paper cites Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding.

Multi-modality Latent Interaction Network for Visual Question Answering Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.544173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.544173Z digest=sha256:13a09343e16348800f6aa4930db229e5846181dd5c1076b4be8b65c0646eed38

Observation 532a131c-4239-4c17-956d-5d8bde91264d · outbound

This paper cites Dy- namic fusion with intra-and inter-modality attention flow for visual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Dy- namic fusion with intra-and inter-modality attention flow for visual question answering

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.492429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.549038Z digest=sha256:0440b1110cd40b20f6282a3f1b4f64c1d749f33e046d58ddaf55badf970febcf

Observation f626cac9-74e9-4d54-8c40-7c3694423d9b · outbound

This paper cites Question-guided hy- brid convolution for visual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Question-guided hy- brid convolution for visual question answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.478287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.553086Z digest=sha256:a64507e2d072c8e682b618911a9f7e1e34d0e0b0e2a20d72750d8e087ef5a47c

Observation 8153b2b5-04b5-446e-a30f-0a7f0c59ab12 · outbound

This paper cites Compact bilinear pooling.

Multi-modality Latent Interaction Network for Visual Question Answering Compact bilinear pooling

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.462487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.556968Z digest=sha256:a7663b1ed2af4b1a6e6f60fa661115e5a7e3c4ae327a83141976857f06bc9cd1

Observation 90176aa6-a73a-4583-9813-cbf714d5130b · outbound

This paper cites 2nd place solution to the gqa challenge.

Multi-modality Latent Interaction Network for Visual Question Answering 2nd place solution to the gqa challenge

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.448325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.561528Z digest=sha256:c2a388324c6529e955192f361c40a9edfa6a5cc729a57b43d6f3746edd6ee62e

Observation d74ce0eb-9090-44eb-9b36-e42a761c56a5 · outbound

This paper cites Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering.

Multi-modality Latent Interaction Network for Visual Question Answering Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.434668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.570021Z digest=sha256:42ccc145916188bbf6a573f6d19dccb6f0237fd2ac5514194b85a125100cfee3

Observation 88e27c44-0f68-48ef-8eb1-1cef736b7665 · outbound

This paper cites Deep residual learning for image recognition.

Multi-modality Latent Interaction Network for Visual Question Answering Deep residual learning for image recognition

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.420596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.573779Z digest=sha256:1ac44cd2c0bd5e85a229996c8ca68ddc0bd4787974f928ac2522841ef72a5fa2

Observation 6a995f6b-c0bd-433d-a828-a4f9984090a5 · outbound

This paper cites Relation networks for object detection.

Multi-modality Latent Interaction Network for Visual Question Answering Relation networks for object detection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.405333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.577727Z digest=sha256:c6927e016f1e9de0d7a7c323713ee46433a6c78c6f00b88563df22c96c1f1c54

Observation 4fcf0b6b-17f1-4c3c-acad-41557ff3f9be · outbound

This paper cites Weakly-supervised Compositional FeatureAggregation for Few-shot Recognition.

Multi-modality Latent Interaction Network for Visual Question Answering Weakly-supervised Compositional FeatureAggregation for Few-shot Recognition

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-14T14:10:16.921728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.581579Z digest=sha256:ba4904cff1265cef957c1df946a378816691b0fdca41fd060b72b36aa794ff0b

Observation 478c626e-c65b-4415-a993-59441475868e · outbound

This paper cites Learning to segment every thing.

Multi-modality Latent Interaction Network for Visual Question Answering Learning to segment every thing

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.392561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.585698Z digest=sha256:f3f31d60b9dc26860fce8cc910f320ea412d74ca77cc12102029839ade08a20c

Observation c800019c-d2df-4c7e-97fd-1769b900edfc · outbound

This paper cites Densely connected convolutional net- works.

Multi-modality Latent Interaction Network for Visual Question Answering Densely connected convolutional net- works

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.589671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.589671Z digest=sha256:ccc389a25cd7572980ca3663eaa95cd4b04c567cf4fa0fdfc2aa0447001fe5f8

Observation 7ffb9b8b-e6e0-4ff7-aa07-2c926e53558b · outbound

This paper cites Video object detection with locally-weighted deformable neighbors.

Multi-modality Latent Interaction Network for Visual Question Answering Video object detection with locally-weighted deformable neighbors

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.372598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.594324Z digest=sha256:7280283ee2632d3263e77efbb11bbc30e43b279cc5159ecd249502a666f8fc23

Observation e35de43b-3009-4694-9a00-0665efae6991 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elemen- tary visual reasoning.

Multi-modality Latent Interaction Network for Visual Question Answering Clevr: A diagnostic dataset for compositional language and elemen- tary visual reasoning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.359045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.598794Z digest=sha256:d8c99cc6682cbb3311f3846fb94fc15335a78211cd7902ae51354993fcc80d72

Observation 7f53e586-a14c-4007-91de-fa7bf28b6dcc · outbound

This paper cites An analysis of visual question answering algorithms.

Multi-modality Latent Interaction Network for Visual Question Answering An analysis of visual question answering algorithms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.340027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.603448Z digest=sha256:0b9a4364592ee8ec3e59b1b73e5ccad1c91fb3db0bce256ad7207ef97dd904b9

Observation c2d20e82-c705-4789-8e7b-4f4a32468a76 · outbound

This paper cites Bilin- ear attention networks.

Multi-modality Latent Interaction Network for Visual Question Answering Bilin- ear attention networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.324081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.607374Z digest=sha256:b0db97c962995c98fc671c41f01f0ca9a9728c203e4f551a1cd5b47e6df73ef5

Observation 789e25c9-c14a-47d9-9448-52a1758c08d1 · outbound

This paper cites Hadamard Product for Low-rank Bilinear Pooling.

Multi-modality Latent Interaction Network for Visual Question Answering Hadamard Product for Low-rank Bilinear Pooling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.611224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.611224Z digest=sha256:2d289bc98784a8926d2b1df663d265cd5d7adeec35e106969095ff3d5905406b

Observation d7099b22-edbf-4701-875c-c95a3db858fa · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Multi-modality Latent Interaction Network for Visual Question Answering Adam: A Method for Stochastic Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.615304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.615304Z digest=sha256:e01cc8ab3c0f3270ebfcb1ba69908a8a6dfc6f7c249d623d8774f02b37e95dc4

Observation fba9b756-99a7-4e13-9aa4-b1a17e9f3783 · outbound

This paper cites Skip-thought vectors.

Multi-modality Latent Interaction Network for Visual Question Answering Skip-thought vectors

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.308746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.621107Z digest=sha256:6e243ee27bc814be068dee7934b83283f77acf9f45acf3742c44ab32a5d7024f

Observation 6cfde9e3-b90f-4126-8fe4-a12cc0d4b3ed · outbound

This paper cites Imagenet classification with deep convolutional neural net- works.

Multi-modality Latent Interaction Network for Visual Question Answering Imagenet classification with deep convolutional neural net- works

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.625287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.625287Z digest=sha256:54d16a9728f6cef795d79b20eab7a0fc216e61fa50678dcecf03dcefe0be667c

Observation 567d933e-59d2-4345-9993-f8dfe5c490fa · outbound

This paper cites Microsoft coco: Common objects in context.

Multi-modality Latent Interaction Network for Visual Question Answering Microsoft coco: Common objects in context

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.629337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.629337Z digest=sha256:0c9f7feb57da7cc0223c261f0a6d663afdd825e33cb341060328bab13b2072f8

Observation 7b111dbf-8d54-49fd-9a3e-3b66434d39ef · outbound

This paper cites Improving referring expression grounding with cross-modal attention-guided erasing.

Multi-modality Latent Interaction Network for Visual Question Answering Improving referring expression grounding with cross-modal attention-guided erasing

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.264602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.633455Z digest=sha256:b11dbffee6c11ada1c53f5f746f9f4f0b3c7246455d92df8985823d62bded42e

Observation fe831b02-d01a-4554-b47e-7eee89b55662 · outbound

This paper cites ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks.

Multi-modality Latent Interaction Network for Visual Question Answering ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.637749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.637749Z digest=sha256:f6321c05334aac0d55f2463819c1dc9eba090cad8dffb2647152a26db45c068c

Observation d1e91f63-aeee-4e88-af30-73411c56a52c · outbound

This paper cites Hierarchical question-image co-attention for visual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Hierarchical question-image co-attention for visual question answering

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.245332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.642418Z digest=sha256:2947f1c6215397ac4b415815dfaa188fd9243b90db481a94795f24dfec5aa069

Observation 25361de9-a941-441c-b375-964800efefb6 · outbound

This paper cites Distributed representations of words and phrases and their compositionality.

Multi-modality Latent Interaction Network for Visual Question Answering Distributed representations of words and phrases and their compositionality

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.228948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.646144Z digest=sha256:60031893347839f49e88744d61abff1818228caabd3ceb49f5c6faefc80b168e

Observation 1a56c1bf-1028-4e01-a1ed-d84b7ac9c4cd · outbound

This paper cites Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.215859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.649799Z digest=sha256:4f1669ff89a3f44e3a0da116d1ae02c8db99e5e98f551e80decd27731717f58e

Observation cea23f11-e8c5-4b1f-a78f-7a5d3a403b19 · outbound

This paper cites Training Recurrent Answering Units with Joint Loss Minimization for VQA.

Multi-modality Latent Interaction Network for Visual Question Answering Training Recurrent Answering Units with Joint Loss Minimization for VQA

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.653933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.653933Z digest=sha256:5a05036333a5d089f823b83c6bb4c3b414b677cffb3cebfd7ee7d6212027babc

Observation 60aacede-dd90-40e1-81e4-0b32d574e339 · outbound

This paper cites Im- age question answering using convolutional neural network with dynamic parameter prediction.

Multi-modality Latent Interaction Network for Visual Question Answering Im- age question answering using convolutional neural network with dynamic parameter prediction

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.201769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.658486Z digest=sha256:868d5a0e2ccee4261fd244f79f9c7926681c06e682416784d8f849079e7281c8

Observation 988bf229-828f-48e0-8335-df53c9ccadca · outbound

This paper cites Learning conditioned graph structures for interpretable vi- sual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Learning conditioned graph structures for interpretable vi- sual question answering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.185787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.664105Z digest=sha256:09d7f66f92252a4d4232b547f9eea3b80f68cd32b86b2109a3949d4bd4d699c9

Observation 9b0fb6fd-9f3c-4f04-9c4f-f11874844548 · outbound

This paper cites Automatic differentiation in pytorch.

Multi-modality Latent Interaction Network for Visual Question Answering Automatic differentiation in pytorch

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.668428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.668428Z digest=sha256:bf321f35ae120630d49ca4c2c9f6a054ed65c153b014549e813a46f2b6939f73

Observation 40a778a7-310d-451a-b131-44197d020028 · outbound

This paper cites Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering.

Multi-modality Latent Interaction Network for Visual Question Answering Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.674020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.674020Z digest=sha256:219dedbbdcd59c60b79df9fca91a538eac588e213c1c556d08cf5a6d1e93bba7

Observation ce22324b-860a-4eae-9220-d29dbf2e4eeb · outbound

This paper cites Glove: Global vectors for word representation.

Multi-modality Latent Interaction Network for Visual Question Answering Glove: Global vectors for word representation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.678574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.678574Z digest=sha256:9c3e4bab80bb4139c310988e74294d14491c856a261a56df6ca49042e59f81db

Observation 68e5d917-76b3-4481-877f-3c6e85b6a2b3 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer.

Multi-modality Latent Interaction Network for Visual Question Answering Film: Visual reasoning with a general conditioning layer

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.146576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.683012Z digest=sha256:e29cbd8dbf6e471ab1bf2c0d58af6facd5f529beb00f01d716fff97e3ba1d792

Observation 7522a3e6-cf18-4dde-80a8-d9cf8eb13b5c · outbound

This paper cites Deep contextualized word representations.

Multi-modality Latent Interaction Network for Visual Question Answering Deep contextualized word representations

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.130481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.688219Z digest=sha256:056d0f4e0bd3b6fb589ebda600adb3367360c5d30301e03752d38947dc70a2de

Observation b48ce64b-0f92-4ee7-a237-1fec3d7f0918 · outbound

This paper cites Language models are unsuper- vised multitask learners.

Multi-modality Latent Interaction Network for Visual Question Answering Language models are unsuper- vised multitask learners

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.692191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.692191Z digest=sha256:f1049b32a0806e724922000e85ad16483a054861763da9e14cac6436b4b80d50

Observation 52362887-7d93-44d1-b0ee-0963074959e1 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

Multi-modality Latent Interaction Network for Visual Question Answering Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.109313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.696716Z digest=sha256:97c942d1b8b35edd588c962dca6944d857e11ecd729dfdf1184ebe458622ee1e

Observation 73f4e08b-e4aa-419a-9b1e-9953aa6767ac · outbound

This paper cites A simple neural network module for relational rea- soning.

Multi-modality Latent Interaction Network for Visual Question Answering A simple neural network module for relational rea- soning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.094891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.701080Z digest=sha256:99b98b1e4f135513f4be9007b4a9856d60b2606ee7fa19ed5e19307df6577b02

Observation 2d6bca56-b516-466c-b929-64a53a53c5bf · outbound

This paper cites Question type guided attention in visual ques- tion answering.

Multi-modality Latent Interaction Network for Visual Question Answering Question type guided attention in visual ques- tion answering

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.081647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.705116Z digest=sha256:f79019174a02f47d6c6154a941a44474481cc3ff911ab4d6cbd84ceda581fc12

Observation 85f2cdf2-0d6e-4cbc-97fe-76967bd6ec6a · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Multi-modality Latent Interaction Network for Visual Question Answering Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.709992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.709992Z digest=sha256:66717f0bf721524dd988e392409c578267aa2e19b3fc78a0acab27d8a0e8177c

Observation b6ee223f-13c4-4ed4-b049-003d9f6bf967 · outbound

This paper cites Attention is all you need.

Multi-modality Latent Interaction Network for Visual Question Answering Attention is all you need

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.069512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.713848Z digest=sha256:00cad2588b005c19748d8cc2fbcf55d441da09b443b057ea75404fd73b3aba68

Observation f9aa4ee0-2de4-4ff1-9ddb-b916780cc722 · outbound

This paper cites Non-local neural networks.

Multi-modality Latent Interaction Network for Visual Question Answering Non-local neural networks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.717747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.717747Z digest=sha256:9b2897a0f5b476ab77b3c751b8bc1c48da14a15fbef648537b432badeb8a8c38

Observation e530c4c5-f0de-477b-8477-f737887aa4dd · outbound

This paper cites Pay Less Attention with Lightweight and Dynamic Convolutions.

Multi-modality Latent Interaction Network for Visual Question Answering Pay Less Attention with Lightweight and Dynamic Convolutions

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.722235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.722235Z digest=sha256:432be4fdf2c35b18fe4ceb0b6ae44bb7dfd1505248e6a7dc1a9c6822e07dc7a7

Observation c1ff441c-aa0b-41ce-b19d-037395704ed9 · outbound

This paper cites Show, attend and tell: Neural image caption gen- eration with visual attention.

Multi-modality Latent Interaction Network for Visual Question Answering Show, attend and tell: Neural image caption gen- eration with visual attention

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.047972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.726466Z digest=sha256:320eba14611ea2403a749aa2d619c1fdce4e4c34f9b377cc0849d0acdaefa0ab

Observation 00a67c6a-a086-41c4-ad6c-cbc161570100 · outbound

This paper cites Stacked attention networks for image question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Stacked attention networks for image question answering

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.034059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.731155Z digest=sha256:7d9840ca128271ccb746ea759e6cfadd8d0dcdea45340995b2ea820bcd49f348

Observation e82286dc-ef8a-4aab-aedb-93690a8ffbb2 · outbound

This paper cites Scene Graph Reasoning with Prior Visual Relationship for Visual Question Answering.

Multi-modality Latent Interaction Network for Visual Question Answering Scene Graph Reasoning with Prior Visual Relationship for Visual Question Answering

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-14T14:10:16.814204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.735502Z digest=sha256:04d704f2455d57d143b6c8c18d014de7af6329e84f77cf739e5f5a5c1f042864

Observation 7ec17cdb-760a-4643-b1c0-c87333ed883f · outbound

This paper cites Explor- ing visual relationship for image captioning.

Multi-modality Latent Interaction Network for Visual Question Answering Explor- ing visual relationship for image captioning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.018843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.739200Z digest=sha256:bcfa7faa0f8a8796e531c02bd2b9b0890b51131d3fb0b10a3ec99eb817a365b0

Observation feb6f77b-7d61-4ecb-b53b-76d295491612 · outbound

This paper cites Beyond bilinear: generalized multimodal factorized high-order pooling for visual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Beyond bilinear: generalized multimodal factorized high-order pooling for visual question answering

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.006135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.743986Z digest=sha256:3cc22aea39f6a0bdd6277bcba98a6d3c5dc73524a37b868d2bf0bff039dc83ff

Observation de7e3c54-6ffb-4a8d-a514-38cbd2ab6e8f · outbound

This paper cites Yin and Yang: Balancing and an- swering binary visual questions.

Multi-modality Latent Interaction Network for Visual Question Answering Yin and Yang: Balancing and an- swering binary visual questions

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:16.991110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.747846Z digest=sha256:b8fd6132c58898dd2fd86f8afce077eef0f6ba8af55af4a017285711d8ee07f3

Observation f4fcaa43-606e-477c-a945-f8fe8977609d · outbound

This paper cites Learning to Count Objects in Natural Images for Visual Question Answering.

Multi-modality Latent Interaction Network for Visual Question Answering Learning to Count Objects in Natural Images for Visual Question Answering

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.755399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.755399Z digest=sha256:c44fa2c4729214458c09719fc53f7eec4400d1bdb937f237f5d6520bd8d2990f

Observation 57c5b259-ddc5-4643-99f5-a6f8b52f7753 · outbound

This paper cites Structured attentions for visual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Structured attentions for visual question answering

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:16.976330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.760619Z digest=sha256:ad58dcf14172229a966d5c63f3f3fdc76305c8d66532d1ef1983c546c6188ca3

Observation abf6c8a0-e384-41c4-a8c4-1eeea82551cd · outbound

This paper cites 2nd Place Solution to the GQA Challenge 2019.

Multi-modality Latent Interaction Network for Visual Question Answering 2nd Place Solution to the GQA Challenge 2019

Reference 2019

Resolution
verified exact
local_arxiv, observed 2026-08-14T14:10:16.939521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.565726Z digest=sha256:1edbf0f8e6eed4c7b8804e7b14cc4251dea505f5b5484f1fd1c5d624e0e058f4

Pith citing papers

Observation 32e70082-8c07-4327-8b3d-e21b107212c7 · inbound

LXMERT: Learning Cross-Modality Encoder Representations from Transformers cites this paper.

LXMERT: Learning Cross-Modality Encoder Representations from Transformers Multi-modality Latent Interaction Network for Visual Question Answering

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-14T12:22:25.833918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-14T12:22:25.530287Z digest=sha256:12c8f5b90e4cf193e08f3dbcba48913bafe8bba6427db97117968db86d92b6e4