Pith. sign in

Paper Citation Record · LEDGER

Multimodal Unified Attention Networks for Vision-and-Language Interactions

As of 15 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 3 inbound Pith citation observations for arXiv:1908.04107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.04107 v2

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T13:55:26.757084Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:22:25.662003Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:39:30.840356Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact1
  • verified fuzzy53
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9783f87-858d-4eb4-812b-65d09ad8df0b · outbound

This paper cites Multimodal deep network embedding with integrated structure and attribute information,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Multimodal deep network embedding with integrated structure and attribute information,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:28.008716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.405870Z digest=sha256:4115e0676cee4d2ee0346db87e28c9f6740aac145608e80f267432cb43e431ff

Observation 70e901c3-cc0d-40bc-a28c-5f92950e3cc7 · outbound

This paper cites Discrim- inative coupled dictionary hashing for fast cross-media retrieval,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Discrim- inative coupled dictionary hashing for fast cross-media retrieval,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.991927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.411450Z digest=sha256:f39b410a7ba4005fbbfb8b3d92801e3b029f7b1dfd49e041d76bcff074e7ee23

Observation f6859418-36ed-429d-b6ab-d99ffa1366d7 · outbound

This paper cites Shared predictive cross-modal deep quantization,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Shared predictive cross-modal deep quantization,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.974963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.416382Z digest=sha256:55c6601098b068613a1144332ae63bc8328b5f7a6f4413177b56bdd2a3114ccc

Observation 8b5b2dbe-a68f-456a-b955-3ba975161532 · outbound

This paper cites Show, attend and tell: Neural image caption generation with visual attention.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Show, attend and tell: Neural image caption generation with visual attention

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.957946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.422009Z digest=sha256:01314c7e2deb0fe174cf2a853e4763c45228c2cab4b9ea10535e0a903ef44e97

Observation 3c17dd8c-eeeb-4498-98a4-fcdeaf39fe57 · outbound

This paper cites From deterministic to generative: multi-modal stochastic rnns for video cap- tioning,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions From deterministic to generative: multi-modal stochastic rnns for video cap- tioning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.940902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.427042Z digest=sha256:a54d69b7a38fd90c1d74ed3c1b972d5eb07ffeb97c3c7c39a9a79d2cd376a70e

Observation 24dfb2cb-350e-48d1-8b1c-330066e7bc33 · outbound

This paper cites Vqa: Visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Vqa: Visual question answering,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.923633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.432661Z digest=sha256:c00d2277a255f021645fbd8c22ace12b7ceece4cef45841c73923e83e105e61d

Observation a712be8a-2deb-440c-8c4c-31ffb95e8ab3 · outbound

This paper cites Ground- ing of textual phrases in images by reconstruction,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Ground- ing of textual phrases in images by reconstruction,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.907818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.439241Z digest=sha256:c15d5594141904c8164863926a7dfe594dde189313d328f2ef274bc3bea9feb9

Observation 8391a958-710e-4be9-8f55-50c76d8493db · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Neural Machine Translation by Jointly Learning to Align and Translate

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.444239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.444239Z digest=sha256:3482f6c045976174926a4960bd2018197564ed11a7b63c39f9ad78f786d4f753

Observation d9c99a69-6c26-4d21-9df4-79d3556906e4 · outbound

This paper cites Recurrent models of visual attention,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Recurrent models of visual attention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.891612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.449809Z digest=sha256:a393b53bdbe2d8aa8d88f285c18a6355bf22d25d3d43a6a63f52b859d645deb2

Observation 76a48858-2880-49ab-904d-fddcbfed84a8 · outbound

This paper cites Draw: A recurrent neural network for image generation,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Draw: A recurrent neural network for image generation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.875785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.454446Z digest=sha256:ed6fc5809791f65f9eeec4ff5dcc18532e4e2297b7ed838e1fcee11c3d0dc804

Observation 025187f2-3715-4aa0-b7cb-2d689b66bb17 · outbound

This paper cites Attention to scale: Scale-aware semantic image segmentation,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Attention to scale: Scale-aware semantic image segmentation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.859785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.459270Z digest=sha256:1f6e69fed6ad15a7c1c196efcf4e328910a8ebb078e30a60239a5dca050e2d3f

Observation cc9ad11a-415f-4503-a2dd-4b836808f174 · outbound

This paper cites Effective Approaches to Attention-based Neural Machine Translation.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Effective Approaches to Attention-based Neural Machine Translation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.465116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.465116Z digest=sha256:0301cce0b4982a99873c2ce087bb7f36286dc7523c7db0fe5a5a9929e317988e

Observation 7b80ff55-9a49-41b6-9d8f-8b089392bd47 · outbound

This paper cites Deep biaffine attention for neural dependency parsing,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Deep biaffine attention for neural dependency parsing,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.842087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.470402Z digest=sha256:32757d1901f67e4d883c17a60d33ed3ed73a92e979f3f10236ccfe0b4866bed6

Observation 83ed30e9-de93-4058-95d8-f2258e795eb8 · outbound

This paper cites A Neural Attention Model for Abstractive Sentence Summarization.

Multimodal Unified Attention Networks for Vision-and-Language Interactions A Neural Attention Model for Abstractive Sentence Summarization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.475737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.475737Z digest=sha256:3f6ac1d302401931c07d9c700ace6293c9b0c5f6c7c59040f1b3f049145f67cf

Observation 5ead648f-1b4b-4cdd-b058-1cc5db435c55 · outbound

This paper cites Stacked attention net- works for image question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Stacked attention net- works for image question answering,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.824687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.481939Z digest=sha256:63ffe26f781c181bbff1332021d47f69f47948cd677a975dbc490749034f6e29

Observation 24d92376-f32e-44b1-9559-3cda4b74118f · outbound

This paper cites Multimodal compact bilinear pooling for visual question answering and visual grounding,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Multimodal compact bilinear pooling for visual question answering and visual grounding,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.807046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.487055Z digest=sha256:d74641db7513d4a91413d99e38e36c41838662cc1035e1ec604597347b9105c6

Observation 93828af5-dd98-46b0-99cb-bde6d74cd5c6 · outbound

This paper cites Hierarchical question-image co-attention for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Hierarchical question-image co-attention for visual question answering,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.790038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.491852Z digest=sha256:b67ab146276368998c890fd03b8ea10a127c117d9061e612b110fada758144a0

Observation 79de54b6-0d38-4847-8c0d-9db88de40266 · outbound

This paper cites Multi-modal factorized bilinear pooling with co-attention learning for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Multi-modal factorized bilinear pooling with co-attention learning for visual question answering,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.771611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.496532Z digest=sha256:2384a916e0dd3f21fdeaff9e4add0334e607a82e9ca2530e869fd400aa506ef6

Observation 2217e021-5f31-4498-b185-6e8a8e8ec933 · outbound

This paper cites Bilinear attention networks,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Bilinear attention networks,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.754499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.501649Z digest=sha256:abb5c29b6e388674173269f57dead8314f6a5fb70e6f63b789cca84ef88b2135

Observation b0fe8842-fcf6-4eea-ba58-794f04ddea9e · outbound

This paper cites Improved fusion of visual and language representations by dense symmetric co-attention for visual question an- swering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Improved fusion of visual and language representations by dense symmetric co-attention for visual question an- swering,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.731604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.506742Z digest=sha256:0752e196a3a6c6a8841d6fb88fea216184379e842ab662c031bb69bd986f246e

Observation 8d40dc6e-7186-4ff5-bb03-b5adeb46d1d9 · outbound

This paper cites Attention is all you need,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Attention is all you need,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.710634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.513064Z digest=sha256:9c7470fd9edf2732eb9d02c99b359c915d52e718a9af0f42312298b4356d43d2

Observation e4dea28e-54d3-44f4-9553-ddbda29ecfb3 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Multimodal Unified Attention Networks for Vision-and-Language Interactions BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.518164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.518164Z digest=sha256:bf8d79eebfa14fb3fcd6202833b231d96e7231604528ec1153ccefd844ea6a41

Observation c50333a2-4c3f-4f51-8dc1-1ac7acbcc3b1 · outbound

This paper cites Non-local neural net- works,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Non-local neural net- works,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.692996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.523097Z digest=sha256:741c3754ee008530f18df601a706f60d81c79f86ccd07f2c58aff47ea28b8eb7

Observation 4d6962c0-7f30-4f83-80b5-ac0d6d0eff6b · outbound

This paper cites Relation networks for object detection,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Relation networks for object detection,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.674066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.527825Z digest=sha256:7ed1d9a95b7dd1e0564438339085a1e443d7c277aa504fac0ac202578ee07d9d

Observation c4fdcde9-6287-4302-8336-8c7a46ddfa72 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Making the v in vqa matter: Elevating the role of image understanding in visual question answering,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.655703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.532554Z digest=sha256:790ca76b47bef1a6a7f1c17728eee606cc13c3f5254725afefdddf1d7e1a10e4

Observation 51ff0f0b-5de1-4252-830d-3b16362c1344 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.638537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.538915Z digest=sha256:2cabfe561e27ce8b303fc2c6125343d50f4c230b557b8e329ad7705cb1d0e5d8

Observation 6e577d9d-4dcb-4137-b9d6-1b817db4de5c · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Referitgame: Referring to objects in photographs of natural scenes,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.543907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.543907Z digest=sha256:2dce5ef93f5eae99e626470ec2e9e4b21636c5fbfa40c8423ed37dd0ea653eaa

Observation f30f20a7-9f63-40b1-b93a-9a8f24d590bd · outbound

This paper cites Generation and comprehension of unambiguous object descriptions,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Generation and comprehension of unambiguous object descriptions,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.610143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.549092Z digest=sha256:b8162f4d6c6d6e91856465822bfb3560c5fab4d49f6c91dc59a75106543bdc81

Observation 6f797fb1-9b33-429f-a636-b0831eafd354 · outbound

This paper cites Simple Baseline for Visual Question Answering.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Simple Baseline for Visual Question Answering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.554516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.554516Z digest=sha256:7e35f6709a7b31468bf53b4965b47e42af8efd4c0ddd4f14f5d02add72e33106

Observation 3d147a9b-af1f-4fb7-82c0-5b22177aadf3 · outbound

This paper cites Hadamard Product for Low-rank Bilinear Pooling,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Hadamard Product for Low-rank Bilinear Pooling,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.592926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.559720Z digest=sha256:2c8ccbe95f7cc0d01ac3dbf5cf91518198cd82e5913f8b636510a35b58199381

Observation 6d9aaaf4-c355-4c10-ab2c-6d3ebe5326e1 · outbound

This paper cites Mutan: Multi- modal tucker fusion for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Mutan: Multi- modal tucker fusion for visual question answering,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.573058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.564802Z digest=sha256:9e43206988685a39034b512595fe2fe6d935137ce3eb5103570dcc2962d5292b

Observation 6d91550f-15d5-4026-88b0-8a0b3148f76a · outbound

This paper cites ABC-CNN: An Attention Based Convolutional Neural Network for Visual Question Answering.

Multimodal Unified Attention Networks for Vision-and-Language Interactions ABC-CNN: An Attention Based Convolutional Neural Network for Visual Question Answering

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.570303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.570303Z digest=sha256:f024ce6faf3254c3048c721d1647d7fe6c6bbff069b09929d2d433cb2d210d56

Observation 109b5f20-38ee-426c-8fee-c9af7d2abde6 · outbound

This paper cites A Focused Dynamic Attention Model for Visual Question Answering.

Multimodal Unified Attention Networks for Vision-and-Language Interactions A Focused Dynamic Attention Model for Visual Question Answering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.575466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.575466Z digest=sha256:d111fa5e7a56c534f5f7e625ed6e3b103018097b71e929c69bf8b56bc38845a1

Observation bb9f24d8-e8e2-4819-8e40-6397ae70a6fa · outbound

This paper cites Where to look: Focus regions for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Where to look: Focus regions for visual question answering,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.553523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.580811Z digest=sha256:bb47a5338c01c18347444e9f587f1e55e317cf7473366c756c0f95b05be6c688

Observation fb59863f-722b-4c6c-be73-b56a7b16b79a · outbound

This paper cites Beyond bilinear: Generalized multi-modal factorized high-order pooling for visual question answer- ing,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Beyond bilinear: Generalized multi-modal factorized high-order pooling for visual question answer- ing,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.533495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.586427Z digest=sha256:ed24aadd1740209bea1d8b2ddbfccb5ebbfcf600254bd474024142eb98c85fd9

Observation 960371a1-dd9e-41eb-aca8-324d0c3487c9 · outbound

This paper cites A joint speaker-listener- reinforcer model for referring expressions,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions A joint speaker-listener- reinforcer model for referring expressions,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.501146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.591097Z digest=sha256:17eea65aa89b599a839062ce4b7e36444f015df3f71ff70c8f4024845dbfde22

Observation fd8e2279-07a3-4bb7-a782-0b2172e20416 · outbound

This paper cites Edge boxes: Locating object proposals from edges,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Edge boxes: Locating object proposals from edges,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.481300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.595646Z digest=sha256:666712854c42196cff9eba78390e4d51fd74d359433f841459c35480d1c97690

Observation eaaee506-d5f9-455d-a353-4f846e9de4db · outbound

This paper cites Deep residual learning for image recognition,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Deep residual learning for image recognition,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.462346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.600632Z digest=sha256:033ad0d8da8c2479161a1857d6c33bd57585561c11d6de98d066668c0b93724e

Observation 848abf49-fe14-48b4-8ed8-c3d6ab67e6a9 · outbound

This paper cites Rethinking diversified and discriminative proposal generation for visual grounding,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Rethinking diversified and discriminative proposal generation for visual grounding,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.444232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.606994Z digest=sha256:afa442c58bc22597e14c93731827c1c8f6528e508d326658029ea7f4f600c548

Observation 09816d98-c73d-422d-96a0-3633ac698948 · outbound

This paper cites Mattnet: Modular attention network for referring expression comprehension,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Mattnet: Modular attention network for referring expression comprehension,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.426916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.611863Z digest=sha256:cfdba826fca902ea47db9be178bc546940737ffb70d57988f0f6128267567e26

Observation 23a1ab66-d516-47ae-a85e-05211370ff20 · outbound

This paper cites Parallel attention: A unified framework for visual object discovery through dialogs and queries,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Parallel attention: A unified framework for visual object discovery through dialogs and queries,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.407914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.616740Z digest=sha256:12e140e868b18db9ea4b66ef20d7e4449ea28fe4a0b0c45559d80ded344b0a3b

Observation 9aee77a4-0b7a-4c8d-8d49-bcb2a792117e · outbound

This paper cites Visual grounding via accumulated attention,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Visual grounding via accumulated attention,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.390683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.622160Z digest=sha256:66d3c929505a6e1f8840c0e5fc98d9031eb81ea1bfb3439c8af470d359b9840e

Observation 619e07f7-7290-4b83-a249-6b0e2840a833 · outbound

This paper cites Be- yond rnns: Positional self-attention with co-attention for video question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Be- yond rnns: Positional self-attention with co-attention for video question answering,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.374092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.627435Z digest=sha256:093e1b639469232114a3d27ee2f805b0e245b7b9cb4b7d503b20c2f226fedd27

Observation b487a113-c177-42c6-920b-73ddb6fbff3b · outbound

This paper cites Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-14T13:55:26.904087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.632542Z digest=sha256:10b06f663c1ee567b551da59bab4ca969396939a475216ab1666e2677632d10a

Observation 8e1ba17d-fc08-41e7-b6a5-609798755313 · outbound

This paper cites Improving language understanding by generative pre-training,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Improving language understanding by generative pre-training,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.357158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.638333Z digest=sha256:49bf1ebfe95223f0e013e14118aaa5a25b5b4b45db90c2dd3f57aea534dec336

Observation 56e1d26b-e8d1-4ce0-84c8-e33d878dfbfc · outbound

This paper cites Factorized bilinear models for image recognition,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Factorized bilinear models for image recognition,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.340667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.643030Z digest=sha256:a74b377b5005307ff449b5bf53513fba857f4a17466cdaff65674529af4b0cb5

Observation 7c91eab1-0e15-4cd8-a276-5bb7285988bc · outbound

This paper cites Layer Normalization.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Layer Normalization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.647648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.647648Z digest=sha256:b851d70bc155c52dbac425822af83f6887f582a2169f4b380be4ea6eceb20ed4

Observation 9c91039a-5194-42a8-a217-dc1fad636ef9 · outbound

This paper cites Glove: Global vectors for word representation.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Glove: Global vectors for word representation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.322774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.652624Z digest=sha256:b2fdec07dc5f6d2800fd94354225c47ef032b8d8b811ec11f51b4eba8b3c1a84

Observation ca5f6cfb-f77c-4d5d-aa6c-22c2b451f2d7 · outbound

This paper cites Long short-term memory,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Long short-term memory,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.657304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.657304Z digest=sha256:f28f0aee388159952572e97ab58982b39f12c2d94f27ab752ec4ca72bb1d9397

Observation e0c6bc30-b4b8-4e0b-a17d-62af13d28368 · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Bottom-up and top-down attention for image captioning and visual question answering,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.290618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.662045Z digest=sha256:01b2f4c4ae02bc1bbaca62d3faf59c980c6cb07406da0b7e0267a78164f07b0b

Observation a01fe0f2-a863-4ebd-9edd-2c4db59909f0 · outbound

This paper cites Tips and Tricks for Visual Question Answering: Learnings from the 2017 Challenge.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Tips and Tricks for Visual Question Answering: Learnings from the 2017 Challenge

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.666553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.666553Z digest=sha256:55d195275f203b8649f95ec5dbddaa61d820691c89939fef73f651142ed74790

Observation 9f73072f-c372-4c77-9609-1f8bc3152ae5 · outbound

This paper cites Fast r-cnn,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Fast r-cnn,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.271577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.671520Z digest=sha256:dfb9217393cf10b63d6b6927b5e5c6fa64febd3f17fef76a7ca9e514aba2a75e

Observation 59e2ed05-d7a5-4aba-9e46-0538c12b7244 · outbound

This paper cites Microsoft coco: Common objects in context,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Microsoft coco: Common objects in context,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.251132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.676051Z digest=sha256:48a33978f8db1c219bf79e829c6c686762690f8651974660324ebaa4b63da590

Observation 5617209c-ca98-4c12-938a-ea8847087550 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Adam: A Method for Stochastic Optimization

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.680958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.680958Z digest=sha256:656b5ed69cef40feb01736ff6c03ff64f336d54f8f6edb8bedcd75a0e0347fd8

Observation ccee27c2-86b0-4da9-b579-2767d2a9f8b1 · outbound

This paper cites Faster r-cnn: Towards real- time object detection with region proposal networks,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Faster r-cnn: Towards real- time object detection with region proposal networks,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.233090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.685644Z digest=sha256:e315f127faebe03a2abdd8dde1291820652ef8b6bd1fbfab356d6df93ca54000

Observation 75b9dd9d-9b4e-436b-a1c0-6d5dae966155 · outbound

This paper cites Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.690283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.690283Z digest=sha256:d91d91656fae79c396fc0e3cf104de991992bda612a6ff8de26d95a1a6658af7

Observation a5d1f2f3-65f2-427c-a4ef-e1246434acb0 · outbound

This paper cites Compositional Attention Networks for Machine Reasoning.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Compositional Attention Networks for Machine Reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.694870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.694870Z digest=sha256:3a36caf1aa6986b83497943312a1312fa73fbcc38d1e4f2acb36017a3166276e

Observation 4b37c205-f5f0-4f0b-87de-189dacaa1a74 · outbound

This paper cites Mask r-cnn,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Mask r-cnn,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.214156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.699662Z digest=sha256:e311026c2ad0ff3cca034b70942a7c402c7621860c0bf91223c7936910c1006f

Observation dc40e05e-a0b8-42f8-8ecb-56e048881c88 · outbound

This paper cites Modeling context in referring expressions,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Modeling context in referring expressions,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.197496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.704689Z digest=sha256:8186b8d92fa2f4cbc8d11208a9d4af996409824e9618acceffaa5848d6460597

Observation 1cfe4aa3-3e50-4d06-89ae-46bce9abedf8 · outbound

This paper cites Learning to count objects in natural images for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Learning to count objects in natural images for visual question answering,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.182018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.709730Z digest=sha256:01302763898d6d1c7ae06c0f95278b85d1a30ca5fae09081e6f579ec2f1240e2

Observation 5513a91a-6151-457e-b88c-d5b26a4e0e03 · outbound

This paper cites Deep modular co- attention networks for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Deep modular co- attention networks for visual question answering,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.166305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.714365Z digest=sha256:2f8518772e168ad7f374be82891959cd804108cbf841091e63535873d65e525a

Observation f5eae359-dab9-4bab-b5de-cd6fe415a587 · outbound

This paper cites Learning to reason: End-to-end module networks for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Learning to reason: End-to-end module networks for visual question answering,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.148630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.719073Z digest=sha256:59d07d525961f7f9600e4b508e36e73854bec1a8783a4b231baa4679ac6bb79c

Observation a20f6b27-33c3-4b36-a99a-bdcd5ada6b5e · outbound

This paper cites A simple neural network module for relational reasoning,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions A simple neural network module for relational reasoning,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.723698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.723698Z digest=sha256:42b39bc9afd3c3a0b2e36ee2b5e8de55068a8b266e4a8632ea979d662ce9e031

Observation 263fe6b5-0431-4cd3-9e01-869ad04fabc7 · outbound

This paper cites Inferring and executing programs for visual reasoning,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Inferring and executing programs for visual reasoning,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.120140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.728584Z digest=sha256:29591fb4b7de36704630bab67943473df299116da51861a9d9b32440dbd64509

Observation 3982b222-2158-42f5-9eaa-2b9aa7fdce41 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Film: Visual reasoning with a general conditioning layer,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.102590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.733087Z digest=sha256:0e0f2e57ec26791384d0d3bc685e86cc35f74e850b1302fb4d08a0a764f6fc9c

Observation fe3807c5-9063-4fe5-b5ee-88a6eb62baa5 · outbound

This paper cites Ssd: Single shot multibox detector,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Ssd: Single shot multibox detector,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.086625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.737651Z digest=sha256:2d96972d7cb2f2817aaefc7dd4b11bc0ed801b079ed3800023d2e87bf15581b9

Observation cb4a4a63-7375-47e7-9e60-4868dba22dee · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.742259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.742259Z digest=sha256:ebc5314dd3ee48d243a99f6c44ef3269e475d7d4286a6ef9f563828c0718c13a

Observation 113d83f9-c6f5-44b9-9627-4a54f7bf7bbf · outbound

This paper cites Referring expression generation and comprehension via attributes,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Referring expression generation and comprehension via attributes,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.069979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.747240Z digest=sha256:9ad3e2b996f7c77d06d1853aad42890ee00d9838889dff6cbf7004f8d61d3497

Observation 056540f1-7b59-45d8-bbe2-4ec13864cf9e · outbound

This paper cites Modeling relationships in referential expressions with compositional modular networks,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Modeling relationships in referential expressions with compositional modular networks,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.052588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.752432Z digest=sha256:154ff037b8b3d7a14445adb1c8035858a40c5348dfc3697bd45c942d774d5e64

Observation ce81e276-226a-4c4f-9822-e6f343ce6cf9 · outbound

This paper cites Grounding referring expressions in images by variational context,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Grounding referring expressions in images by variational context,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.036101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.757084Z digest=sha256:608eca17f92c37027972399b1d6038ca8f193a275c99d3c6b8fe28ddadd7c842

Pith citing papers

Observation e801b09e-8839-475b-bff5-3d70b5134379 · inbound

LXMERT: Learning Cross-Modality Encoder Representations from Transformers cites this paper.

LXMERT: Learning Cross-Modality Encoder Representations from Transformers Multimodal Unified Attention Networks for Vision-and-Language Interactions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-14T12:22:25.662003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T12:22:25.662003Z digest=sha256:666f1c9715bcab20922812614c40513c5a179c3e9aab48b8f6cc03f557376984

Observation 727d9b13-5265-4449-ad2c-613515bb85a3 · inbound

ViASNet: A Video Ad Saliency Network for Predicting Dynamic Saliency and Viewer Engagement cites this paper.

ViASNet: A Video Ad Saliency Network for Predicting Dynamic Saliency and Viewer Engagement Multimodal Unified Attention Networks for Vision-and-Language Interactions

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:53:15.881026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T08:50:03.872247Z digest=sha256:1eae737f3b769f4219c71e5f9ad5462be3cfcbd2aa9f43eb048c163eb97c2ad1

Observation 419bf000-2dcf-42f0-a8e1-6427010faba4 · inbound

Alzheimer's Disease Diagnosis using a Multimodal Approach with 3D MRI and PET cites this paper.

Alzheimer's Disease Diagnosis using a Multimodal Approach with 3D MRI and PET Multimodal Unified Attention Networks for Vision-and-Language Interactions

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:39:30.843132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T17:44:50.842397Z digest=sha256:bb4e0fe1f4d0995950f3b9ac8273656a78205b5cbdcea57576f3313360fa997b