Pith. sign in

Paper Citation Record · LEDGER

Multimodal Unified Attention Networks for Vision-and-Language Interactions

As of 16 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 3 inbound Pith citation observations for arXiv:1908.04107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.04107 v2

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T13:55:26.757084Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:22:25.662003Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:39:30.840356Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact1
  • verified fuzzy53
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9783f87-858d-4eb4-812b-65d09ad8df0b · outbound

This paper cites Multimodal deep network embedding with integrated structure and attribute information,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Multimodal deep network embedding with integrated structure and attribute information,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:28.008716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.405870Z digest=sha256:a57563f11e8e253c2a8f16d4b24e5c0c2ab1c4bf7e06ed3b3260337eaa65eb18

Observation 70e901c3-cc0d-40bc-a28c-5f92950e3cc7 · outbound

This paper cites Discrim- inative coupled dictionary hashing for fast cross-media retrieval,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Discrim- inative coupled dictionary hashing for fast cross-media retrieval,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.991927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.411450Z digest=sha256:a8bdec09f3719bb4777358e1b0b37aeffeffe8f2b34e01740987d1692112d423

Observation f6859418-36ed-429d-b6ab-d99ffa1366d7 · outbound

This paper cites Shared predictive cross-modal deep quantization,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Shared predictive cross-modal deep quantization,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.974963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.416382Z digest=sha256:fbcfb75ca15c5a1bdd8dd5bb317a7c6bcd827c638276ef48dc9480c4acbb8c8d

Observation 8b5b2dbe-a68f-456a-b955-3ba975161532 · outbound

This paper cites Show, attend and tell: Neural image caption generation with visual attention.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Show, attend and tell: Neural image caption generation with visual attention

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.957946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.422009Z digest=sha256:9abca04ffd7c9f55f8649f74555ce6fcc5337da24a5f89c9b2cb0108bd4f0086

Observation 3c17dd8c-eeeb-4498-98a4-fcdeaf39fe57 · outbound

This paper cites From deterministic to generative: multi-modal stochastic rnns for video cap- tioning,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions From deterministic to generative: multi-modal stochastic rnns for video cap- tioning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.940902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.427042Z digest=sha256:4253f770c334fc4a25bfa058ee5362965d54add49f930a56e0ff45ebd3c1f891

Observation 24dfb2cb-350e-48d1-8b1c-330066e7bc33 · outbound

This paper cites Vqa: Visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Vqa: Visual question answering,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.923633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.432661Z digest=sha256:b5b3177ea2bcb899192ce947e8fe94ddf3fcb89047c1fa4595c51534fae778fa

Observation a712be8a-2deb-440c-8c4c-31ffb95e8ab3 · outbound

This paper cites Ground- ing of textual phrases in images by reconstruction,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Ground- ing of textual phrases in images by reconstruction,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.907818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.439241Z digest=sha256:217dc0dbdb106f4e3e0979695bdb1a35c7a38cd426b32033f365bdd5433ddbcd

Observation 8391a958-710e-4be9-8f55-50c76d8493db · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Neural Machine Translation by Jointly Learning to Align and Translate

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.444239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.444239Z digest=sha256:2465ce23426af7e53a6082e93e5a28c48b423ac107b3d89892e07739859f5312

Observation d9c99a69-6c26-4d21-9df4-79d3556906e4 · outbound

This paper cites Recurrent models of visual attention,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Recurrent models of visual attention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.891612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.449809Z digest=sha256:d1a8f756853f7a85ff7e32641cd2e2ebfa5f2f35b9cfe8ec2072ee708e0f9fd9

Observation 76a48858-2880-49ab-904d-fddcbfed84a8 · outbound

This paper cites Draw: A recurrent neural network for image generation,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Draw: A recurrent neural network for image generation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.875785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.454446Z digest=sha256:e3622ae65a944ffdbdb24dbcdd6ec27b784257547bdc6b187b98a92ed779fff8

Observation 025187f2-3715-4aa0-b7cb-2d689b66bb17 · outbound

This paper cites Attention to scale: Scale-aware semantic image segmentation,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Attention to scale: Scale-aware semantic image segmentation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.859785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.459270Z digest=sha256:5673e6c4571ab9abf7c1fce1ab88531fb7cd2eec627737e84695583d801f754a

Observation cc9ad11a-415f-4503-a2dd-4b836808f174 · outbound

This paper cites Effective Approaches to Attention-based Neural Machine Translation.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Effective Approaches to Attention-based Neural Machine Translation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.465116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.465116Z digest=sha256:d1d274c89fbc5e42fcbed20ff7561a7a9fd04cfb1b6748bc17448d8a25d63304

Observation 7b80ff55-9a49-41b6-9d8f-8b089392bd47 · outbound

This paper cites Deep biaffine attention for neural dependency parsing,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Deep biaffine attention for neural dependency parsing,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.842087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.470402Z digest=sha256:5b328d326165a4ff9e42737e4b826aa223fb37056a2d2a91a0c9f36c13ba7a05

Observation 83ed30e9-de93-4058-95d8-f2258e795eb8 · outbound

This paper cites A Neural Attention Model for Abstractive Sentence Summarization.

Multimodal Unified Attention Networks for Vision-and-Language Interactions A Neural Attention Model for Abstractive Sentence Summarization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.475737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.475737Z digest=sha256:23e6de7b566a433e9f5b8ab2fe9f6a673c58e07b2c34301eaaef671e2fb71bc7

Observation 5ead648f-1b4b-4cdd-b058-1cc5db435c55 · outbound

This paper cites Stacked attention net- works for image question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Stacked attention net- works for image question answering,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.824687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.481939Z digest=sha256:0d86aaf40fa4e1bf97ba08f1c37f72568e367bb26025adae9200cda46d93613c

Observation 24d92376-f32e-44b1-9559-3cda4b74118f · outbound

This paper cites Multimodal compact bilinear pooling for visual question answering and visual grounding,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Multimodal compact bilinear pooling for visual question answering and visual grounding,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.807046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.487055Z digest=sha256:f2ec3e61a2fb69513c424b95ee3b1c753925dcb8f88e279436acca756ff71076

Observation 93828af5-dd98-46b0-99cb-bde6d74cd5c6 · outbound

This paper cites Hierarchical question-image co-attention for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Hierarchical question-image co-attention for visual question answering,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.790038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.491852Z digest=sha256:baeda6bbb53af332016dc1a4e72180cf8a9bb4ab408645c2565679a409b78360

Observation 79de54b6-0d38-4847-8c0d-9db88de40266 · outbound

This paper cites Multi-modal factorized bilinear pooling with co-attention learning for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Multi-modal factorized bilinear pooling with co-attention learning for visual question answering,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.771611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.496532Z digest=sha256:e3461fe9ad07a61b22f8bae375561ceb5e404a7e9ada049d99a057a5b6aebc62

Observation 2217e021-5f31-4498-b185-6e8a8e8ec933 · outbound

This paper cites Bilinear attention networks,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Bilinear attention networks,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.754499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.501649Z digest=sha256:4328ff10f43eace3802c9ab55f02a6873223e5048c02a2e62502ba5db9364bf7

Observation b0fe8842-fcf6-4eea-ba58-794f04ddea9e · outbound

This paper cites Improved fusion of visual and language representations by dense symmetric co-attention for visual question an- swering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Improved fusion of visual and language representations by dense symmetric co-attention for visual question an- swering,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.731604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.506742Z digest=sha256:df7490378248316010c3eec30cded00a09cd498a40add6b0383db68f19735a5e

Observation 8d40dc6e-7186-4ff5-bb03-b5adeb46d1d9 · outbound

This paper cites Attention is all you need,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Attention is all you need,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.710634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.513064Z digest=sha256:e73457bbc3d039143c99a1c4b223c733c7033347ed6684e3c08cb52637961971

Observation e4dea28e-54d3-44f4-9553-ddbda29ecfb3 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Multimodal Unified Attention Networks for Vision-and-Language Interactions BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.518164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.518164Z digest=sha256:5ea3025ef2e4c647a7bff7ae340b025cefb37fd47cfe229118a08a97eae98e3e

Observation c50333a2-4c3f-4f51-8dc1-1ac7acbcc3b1 · outbound

This paper cites Non-local neural net- works,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Non-local neural net- works,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.692996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.523097Z digest=sha256:34cd340c15f089e9c9798c599921047002a084140e0408dc087e82a5f2be56b7

Observation 4d6962c0-7f30-4f83-80b5-ac0d6d0eff6b · outbound

This paper cites Relation networks for object detection,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Relation networks for object detection,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.674066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.527825Z digest=sha256:c7210bc9f3f7016582bc70113afee4be5262640c5a431c1ccfc5d3c6257262ee

Observation c4fdcde9-6287-4302-8336-8c7a46ddfa72 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Making the v in vqa matter: Elevating the role of image understanding in visual question answering,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.655703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.532554Z digest=sha256:fbbafd5994d987e7996b0b8931b12ff37faca339addca3cf14514cb40231bbf7

Observation 51ff0f0b-5de1-4252-830d-3b16362c1344 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.638537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.538915Z digest=sha256:6198d51ce4fa73537861c92286be28e7aa15be7afec7dcf224692d499964102d

Observation 6e577d9d-4dcb-4137-b9d6-1b817db4de5c · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Referitgame: Referring to objects in photographs of natural scenes,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.543907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.543907Z digest=sha256:451b7853a14b84f72fae9462bd71aa90adad864f5b741ee38d631c70640c122d

Observation f30f20a7-9f63-40b1-b93a-9a8f24d590bd · outbound

This paper cites Generation and comprehension of unambiguous object descriptions,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Generation and comprehension of unambiguous object descriptions,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.610143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.549092Z digest=sha256:4c698a6fe027a4e3cca009bb945795be9a4181536b82f826656241f923147e98

Observation 6f797fb1-9b33-429f-a636-b0831eafd354 · outbound

This paper cites Simple Baseline for Visual Question Answering.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Simple Baseline for Visual Question Answering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.554516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.554516Z digest=sha256:78252ca932620a50d6cb49d63ba558bce48131fea76ebdd076d417d3a08d2b3d

Observation 3d147a9b-af1f-4fb7-82c0-5b22177aadf3 · outbound

This paper cites Hadamard Product for Low-rank Bilinear Pooling,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Hadamard Product for Low-rank Bilinear Pooling,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.592926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.559720Z digest=sha256:aa9a02af928a4dca5fc1e870e997ab604f9aa5da00fe0bb8ddddbae2e5afa83e

Observation 6d9aaaf4-c355-4c10-ab2c-6d3ebe5326e1 · outbound

This paper cites Mutan: Multi- modal tucker fusion for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Mutan: Multi- modal tucker fusion for visual question answering,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.573058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.564802Z digest=sha256:3944ec7997caca5eb4a30f60a07d820bc8e5101e9dce27ebd87c4600f473915e

Observation 6d91550f-15d5-4026-88b0-8a0b3148f76a · outbound

This paper cites ABC-CNN: An Attention Based Convolutional Neural Network for Visual Question Answering.

Multimodal Unified Attention Networks for Vision-and-Language Interactions ABC-CNN: An Attention Based Convolutional Neural Network for Visual Question Answering

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.570303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.570303Z digest=sha256:80487fbd53d6107c36c0eb742e350c9611d1acd2345fb95ca497f9a49424ec37

Observation 109b5f20-38ee-426c-8fee-c9af7d2abde6 · outbound

This paper cites A Focused Dynamic Attention Model for Visual Question Answering.

Multimodal Unified Attention Networks for Vision-and-Language Interactions A Focused Dynamic Attention Model for Visual Question Answering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.575466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.575466Z digest=sha256:5dd530e9706f6767e814ae1e92ed42719677f03886f399325471c1466926527b

Observation bb9f24d8-e8e2-4819-8e40-6397ae70a6fa · outbound

This paper cites Where to look: Focus regions for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Where to look: Focus regions for visual question answering,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.553523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.580811Z digest=sha256:8dd2625be2216fcea3de510749fffb107108edd4f9e03e227f8e8951af74679c

Observation fb59863f-722b-4c6c-be73-b56a7b16b79a · outbound

This paper cites Beyond bilinear: Generalized multi-modal factorized high-order pooling for visual question answer- ing,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Beyond bilinear: Generalized multi-modal factorized high-order pooling for visual question answer- ing,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.533495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.586427Z digest=sha256:60b1e88bee5027d1bdeed086d60c6161b189b3fd1827a58cecc35e8b449481f4

Observation 960371a1-dd9e-41eb-aca8-324d0c3487c9 · outbound

This paper cites A joint speaker-listener- reinforcer model for referring expressions,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions A joint speaker-listener- reinforcer model for referring expressions,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.501146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.591097Z digest=sha256:8223f8ba86c2eec038a65e3102162c3fa32425ee3869cd76bcd7ba71f21f8b58

Observation fd8e2279-07a3-4bb7-a782-0b2172e20416 · outbound

This paper cites Edge boxes: Locating object proposals from edges,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Edge boxes: Locating object proposals from edges,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.481300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.595646Z digest=sha256:80c861173336ce5505ce6ef9867179a3aca3a09ae289dd47e644c55dff1a14cd

Observation eaaee506-d5f9-455d-a353-4f846e9de4db · outbound

This paper cites Deep residual learning for image recognition,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Deep residual learning for image recognition,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.462346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.600632Z digest=sha256:bdfb2264ec41bba71f515a500982d87e236f5d53b9e6a462ca8e4d5365ffbcd6

Observation 848abf49-fe14-48b4-8ed8-c3d6ab67e6a9 · outbound

This paper cites Rethinking diversified and discriminative proposal generation for visual grounding,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Rethinking diversified and discriminative proposal generation for visual grounding,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.444232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.606994Z digest=sha256:5eeab319a50d97b8a9ee057fa0c00423c1b00bac8a95f6d04d4471200dd09bde

Observation 09816d98-c73d-422d-96a0-3633ac698948 · outbound

This paper cites Mattnet: Modular attention network for referring expression comprehension,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Mattnet: Modular attention network for referring expression comprehension,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.426916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.611863Z digest=sha256:8a4fc75b27df3defb5573cf136c8dbf1a11d1d3ef47f3945eab4f092e6fc0d8b

Observation 23a1ab66-d516-47ae-a85e-05211370ff20 · outbound

This paper cites Parallel attention: A unified framework for visual object discovery through dialogs and queries,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Parallel attention: A unified framework for visual object discovery through dialogs and queries,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.407914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.616740Z digest=sha256:89831754dacaf197cad2395c57b42afde12b0e2267ea472da95cda721e145ae8

Observation 9aee77a4-0b7a-4c8d-8d49-bcb2a792117e · outbound

This paper cites Visual grounding via accumulated attention,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Visual grounding via accumulated attention,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.390683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.622160Z digest=sha256:b0d3c17cee26aea5b2a9d447e31f9d620c8138144a61b50e36bc4256b18ea9c4

Observation 619e07f7-7290-4b83-a249-6b0e2840a833 · outbound

This paper cites Be- yond rnns: Positional self-attention with co-attention for video question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Be- yond rnns: Positional self-attention with co-attention for video question answering,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.374092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.627435Z digest=sha256:7a46eebdc9d71fc0d02a2d3e0e36486084e8a68290ac7ae3318c8f7ee43c5c7c

Observation b487a113-c177-42c6-920b-73ddb6fbff3b · outbound

This paper cites Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-14T13:55:26.904087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.632542Z digest=sha256:86ec9d0bad0c4457e69efa78d83a2c58b7a3a576d1c4658b7f0d65cd266f0d36

Observation 8e1ba17d-fc08-41e7-b6a5-609798755313 · outbound

This paper cites Improving language understanding by generative pre-training,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Improving language understanding by generative pre-training,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.357158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.638333Z digest=sha256:7b0c18b7d8755ecf289efa88714ad9b1f48c7ad92352eaf25082d97864c70992

Observation 56e1d26b-e8d1-4ce0-84c8-e33d878dfbfc · outbound

This paper cites Factorized bilinear models for image recognition,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Factorized bilinear models for image recognition,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.340667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.643030Z digest=sha256:4d324f301ea46627804ce9de0cf21e514844335468336c4b1aebd1fab0c6a333

Observation 7c91eab1-0e15-4cd8-a276-5bb7285988bc · outbound

This paper cites Layer Normalization.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Layer Normalization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.647648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.647648Z digest=sha256:417e9ffc9d1ffd2bc88b70d31f93cd146a18f594b05f26daca5bb995565ea543

Observation 9c91039a-5194-42a8-a217-dc1fad636ef9 · outbound

This paper cites Glove: Global vectors for word representation.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Glove: Global vectors for word representation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.322774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.652624Z digest=sha256:28f4a7220e3b0b93242166c3c703759ef0ba1f90eed4bdeb31c4a1e23178ef6e

Observation ca5f6cfb-f77c-4d5d-aa6c-22c2b451f2d7 · outbound

This paper cites Long short-term memory,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Long short-term memory,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.657304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.657304Z digest=sha256:e1cff225aa4beb1d9498b3450ed5f7abfb2003dbb6521da9a6761ff26dfc6c89

Observation e0c6bc30-b4b8-4e0b-a17d-62af13d28368 · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Bottom-up and top-down attention for image captioning and visual question answering,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.290618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.662045Z digest=sha256:da4ec757bee9c359b1a81e7cb8585d3aaa3fa5aa1d4bbe1b98589fd2c83151ff

Observation a01fe0f2-a863-4ebd-9edd-2c4db59909f0 · outbound

This paper cites Tips and Tricks for Visual Question Answering: Learnings from the 2017 Challenge.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Tips and Tricks for Visual Question Answering: Learnings from the 2017 Challenge

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.666553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.666553Z digest=sha256:7b2aacaf212fc40587814aa6455b39cf76bd707c34e09387db20e37a8b8e4e78

Observation 9f73072f-c372-4c77-9609-1f8bc3152ae5 · outbound

This paper cites Fast r-cnn,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Fast r-cnn,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.271577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.671520Z digest=sha256:6e3184c246d74103d62e00f647c758a1b43a230258a272a37085229425c62470

Observation 59e2ed05-d7a5-4aba-9e46-0538c12b7244 · outbound

This paper cites Microsoft coco: Common objects in context,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Microsoft coco: Common objects in context,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.251132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.676051Z digest=sha256:063a268f180378fc72c3c8a6ade4f7683c53f2301f8978adf6139e3c9e854595

Observation 5617209c-ca98-4c12-938a-ea8847087550 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Adam: A Method for Stochastic Optimization

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.680958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.680958Z digest=sha256:c225f60499919ea5717a424b0ac33f5ec6d383efc280ecb0f293560b06f1f4ce

Observation ccee27c2-86b0-4da9-b579-2767d2a9f8b1 · outbound

This paper cites Faster r-cnn: Towards real- time object detection with region proposal networks,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Faster r-cnn: Towards real- time object detection with region proposal networks,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.233090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.685644Z digest=sha256:a82e3444b507c939b286702d625680899c77aaa41b1429f0802dcb3503c12c26

Observation 75b9dd9d-9b4e-436b-a1c0-6d5dae966155 · outbound

This paper cites Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.690283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.690283Z digest=sha256:fb6993560cb9a512da9f44be65d34665c0aff12280d16d8809f67f5ae85af30f

Observation a5d1f2f3-65f2-427c-a4ef-e1246434acb0 · outbound

This paper cites Compositional Attention Networks for Machine Reasoning.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Compositional Attention Networks for Machine Reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.694870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.694870Z digest=sha256:ac23bbf2e9b30aa0aa11e60606c37edcaa641a260af41a34e9951566a88af7bc

Observation 4b37c205-f5f0-4f0b-87de-189dacaa1a74 · outbound

This paper cites Mask r-cnn,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Mask r-cnn,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.214156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.699662Z digest=sha256:0e4584330a21198dafb245ee0554e8f55ed090b7d111a2bb223c40dd32d1932c

Observation dc40e05e-a0b8-42f8-8ecb-56e048881c88 · outbound

This paper cites Modeling context in referring expressions,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Modeling context in referring expressions,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.197496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.704689Z digest=sha256:8c2c00801aa8f2b75bfbc93e8752bd7a226aa1984b85e27fcf1a8a6da512c0a1

Observation 1cfe4aa3-3e50-4d06-89ae-46bce9abedf8 · outbound

This paper cites Learning to count objects in natural images for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Learning to count objects in natural images for visual question answering,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.182018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.709730Z digest=sha256:b734f9671db5264d967885f40ed85b890b2c20d585d1ccd4c532604252af5142

Observation 5513a91a-6151-457e-b88c-d5b26a4e0e03 · outbound

This paper cites Deep modular co- attention networks for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Deep modular co- attention networks for visual question answering,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.166305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.714365Z digest=sha256:e3ff5f469ed037da6fb38e7bfcd1ac9bdad694a46fe55f3cac7c2c15857176ee

Observation f5eae359-dab9-4bab-b5de-cd6fe415a587 · outbound

This paper cites Learning to reason: End-to-end module networks for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Learning to reason: End-to-end module networks for visual question answering,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.148630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.719073Z digest=sha256:cbc79b5d91e2a65974a30b858c2147aa4ea57758ca64aaa4740452a2da45f2e7

Observation a20f6b27-33c3-4b36-a99a-bdcd5ada6b5e · outbound

This paper cites A simple neural network module for relational reasoning,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions A simple neural network module for relational reasoning,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.723698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.723698Z digest=sha256:009bdeec43511d190fe85ed60c4eb476907c3b9f3a2eccd44740bd53423967df

Observation 263fe6b5-0431-4cd3-9e01-869ad04fabc7 · outbound

This paper cites Inferring and executing programs for visual reasoning,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Inferring and executing programs for visual reasoning,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.120140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.728584Z digest=sha256:2e113d83fdbbc24c6669c87a0275273721afaf97736360bb2fedbf081789aed8

Observation 3982b222-2158-42f5-9eaa-2b9aa7fdce41 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Film: Visual reasoning with a general conditioning layer,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.102590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.733087Z digest=sha256:f1099f27117f1f4d364a36f3c80651b617ea14643c664d7a8a9fc27edad6807e

Observation fe3807c5-9063-4fe5-b5ee-88a6eb62baa5 · outbound

This paper cites Ssd: Single shot multibox detector,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Ssd: Single shot multibox detector,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.086625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.737651Z digest=sha256:d1e8aee16ee520fc1675e8a26a172d572df7a920f6c09e864573a4de7e53fd6b

Observation cb4a4a63-7375-47e7-9e60-4868dba22dee · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.742259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.742259Z digest=sha256:e592573abd09c71af8f0451d8556d0cffb7afb3d1e3ebd08d146da0eee04d92a

Observation 113d83f9-c6f5-44b9-9627-4a54f7bf7bbf · outbound

This paper cites Referring expression generation and comprehension via attributes,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Referring expression generation and comprehension via attributes,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.069979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.747240Z digest=sha256:337d718fa183196fe8030bc53f6dbb8b77632282a9eb2f655716ded6587559ed

Observation 056540f1-7b59-45d8-bbe2-4ec13864cf9e · outbound

This paper cites Modeling relationships in referential expressions with compositional modular networks,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Modeling relationships in referential expressions with compositional modular networks,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.052588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.752432Z digest=sha256:42b545cc407549b5d5d1b69058b04093324436aa36177765c41cdcdd3ebe4150

Observation ce81e276-226a-4c4f-9822-e6f343ce6cf9 · outbound

This paper cites Grounding referring expressions in images by variational context,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Grounding referring expressions in images by variational context,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.036101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T13:55:26.757084Z digest=sha256:23585b6bbe722df1dd3d752ec77e66a81488524b43305ab4071929a5527dbf33

Pith citing papers

Observation e801b09e-8839-475b-bff5-3d70b5134379 · inbound

LXMERT: Learning Cross-Modality Encoder Representations from Transformers cites this paper.

LXMERT: Learning Cross-Modality Encoder Representations from Transformers Multimodal Unified Attention Networks for Vision-and-Language Interactions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-14T12:22:25.662003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T12:22:25.662003Z digest=sha256:fa524faee19240b12dcc7112ef46d2d4fa7404cdc606786dfd029fb7d4a8b1f9

Observation 727d9b13-5265-4449-ad2c-613515bb85a3 · inbound

ViASNet: A Video Ad Saliency Network for Predicting Dynamic Saliency and Viewer Engagement cites this paper.

ViASNet: A Video Ad Saliency Network for Predicting Dynamic Saliency and Viewer Engagement Multimodal Unified Attention Networks for Vision-and-Language Interactions

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:53:15.881026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T08:50:03.872247Z digest=sha256:f8c20240f1468a528cde3757acd603bbc920937951d6d616b058971429fe6794

Observation 419bf000-2dcf-42f0-a8e1-6427010faba4 · inbound

Alzheimer's Disease Diagnosis using a Multimodal Approach with 3D MRI and PET cites this paper.

Alzheimer's Disease Diagnosis using a Multimodal Approach with 3D MRI and PET Multimodal Unified Attention Networks for Vision-and-Language Interactions

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:39:30.843132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T17:44:50.842397Z digest=sha256:064f9865b3503e494f0fdb249fc387b69d83fda8b439e0bbb53e13d068e3643b