Pith. sign in

Paper Citation Record · LEDGER

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering

As of 15 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:1908.04950.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.04950 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T13:32:25.742010Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb66dcb0-3718-4272-97f1-4862dde0b7a5 · outbound

This paper cites Blindfold Baselines for Embodied QA.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Blindfold Baselines for Embodied QA

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.392507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.392507Z digest=sha256:b442df62118ded7d0523c540fbe303bdc976e26ad2098607565fb63f9f59423c

Observation 3ceb717c-cb95-4285-9818-7c304bbd45d0 · outbound

This paper cites VQA: Visual question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering VQA: Visual question answering

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:27.161905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.403046Z digest=sha256:fa4b90c2079106acb46a74e4686ced094e2566ddfe099cdda5aa5592bf78427c

Observation bb7cb70d-7249-41d5-b63f-432c24679e83 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Neural Machine Translation by Jointly Learning to Align and Translate

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.416278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.416278Z digest=sha256:1f9d6c56e8f21aed46ec8bd8761e781146dce48bc70322a0a8a00e251980759d

Observation f698ce84-25f4-4e73-a082-099eaa098921 · outbound

This paper cites Systematic Generalization: What Is Required and Can It Be Learned?.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Systematic Generalization: What Is Required and Can It Be Learned?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.423479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.423479Z digest=sha256:a97e44279fe7a2dc9315197ce1fc0cd3e3da62e9bb83a6a9a405510b20af48eb

Observation d9435600-09ec-4fb2-b0af-0f1bf2f7a7e5 · outbound

This paper cites MUREL: Multimodal Relational Reasoning for Visual Question Answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering MUREL: Multimodal Relational Reasoning for Visual Question Answering

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:27.133297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.432553Z digest=sha256:56cc7df8dace2f317eb09c4a67bbc009dbfcccaeeffeb67b2d32fa7428963a96

Observation 21037c04-3690-42e4-9a09-ab511b70b655 · outbound

This paper cites Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.441540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.441540Z digest=sha256:d69b162c6475c8e7bc952381653931cc77952969d2b90c3450f4f79258e4e185

Observation d902790f-c848-4070-b491-26252b6a3801 · outbound

This paper cites Embodied Question Answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Embodied Question Answering

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:27.106821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.449561Z digest=sha256:fbb954d6ccc7a284e464bac84389cfed58961de2a5a76ea7c9995c226ab29080

Observation 0611e946-1323-42c6-b858-d789e6676321 · outbound

This paper cites Neural Modular Control for Embodied Question Answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Neural Modular Control for Embodied Question Answering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.457445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.457445Z digest=sha256:9b28d85616a9ec0902bbf5ce3454d1ef7e4d5be177ec7f66d98ca69d9a3c218f

Observation e7281b40-acab-49ad-ad69-0787fb0744d3 · outbound

This paper cites Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.464101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.464101Z digest=sha256:51d3751c7110046107562aea509e00a8dbcbcde5d65f20a846dbc53362ed7248

Observation c4b69e52-cbce-46cf-99b2-be8ad2726443 · outbound

This paper cites Deep sparse rectifier neural net- works.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Deep sparse rectifier neural net- works

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:27.084683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.469712Z digest=sha256:132870e306b39923c500eba6a7915ffdf3f724dc6962b16a7ea8fec57889e242

Observation 25ae027a-3402-45cb-bb46-353cc3a5dae2 · outbound

This paper cites IQA: Visual question answering in interactive environments.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering IQA: Visual question answering in interactive environments

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:27.054516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.475855Z digest=sha256:81958b37a209bf7e079440db668cddb41b722a94a8c7316fa8c947f7fe580612

Observation 580fc635-eca8-4468-a0ab-dde7ab6ca1f2 · outbound

This paper cites Long short-term memory.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Long short-term memory

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.481520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.481520Z digest=sha256:552c633eabfd04fcc6acd3a246516fc70bd256e2bc2b95467cbc2403b8444f18

Observation 4c4f30e5-db58-4caf-9271-0de752a17aba · outbound

This paper cites Learning to reason: End-to-end module networks for visual question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning to reason: End-to-end module networks for visual question answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:27.007765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.489048Z digest=sha256:b415601c9eb5a220b53c834cef32a2dfe387eff25e9b5433546477832467090a

Observation 8668e006-e0f2-42ff-a27d-75aa378c41fd · outbound

This paper cites Compositional Attention Networks for Machine Reasoning.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Compositional Attention Networks for Machine Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.498398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.498398Z digest=sha256:0e38ee88d553e95c22a64e64fac3a43a254899bfa1403e7204b7160f23b8866e

Observation 5d8ca0f3-662c-475c-9429-cf4b9f0a345d · outbound

This paper cites GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.507069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.507069Z digest=sha256:61f0c2faea43db244083cc37a932df1871b16a050ba24119655da5fda2638088

Observation 2f4ecd35-991c-4f79-838f-d2be5c8057e6 · outbound

This paper cites Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.514240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.514240Z digest=sha256:6ae465fb2d4e212138b45db87aaef888cdf6115bb285c6522d273f9ba5ddeb6d

Observation aac89881-1bbe-4c80-97d8-1b4d1b361ced · outbound

This paper cites CLEVR: A diagnostic dataset for compo- sitional language and elementary visual reasoning.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering CLEVR: A diagnostic dataset for compo- sitional language and elementary visual reasoning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.983480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.524485Z digest=sha256:2216acfc072fe4541b5ee194a10e5ec7ee854013d88dabb97cc49afba7ac79be

Observation 5667ffb2-fe37-4b0f-9b30-5dedac342b26 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Adam: A Method for Stochastic Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.531651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.531651Z digest=sha256:635fb3441d933d679c963af9066bdca81f6c0d8129c5498f9d75ab30220fdaf5

Observation 8c93cb0d-202b-46c5-a396-3ae3e9b56257 · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.957511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.544151Z digest=sha256:7f288389e4adb895d6ec505c190692e4d0192b47bbd4fd4c9d6b95de05125ac0

Observation a5d61e02-1473-4008-9229-22968a93c996 · outbound

This paper cites TVQA: Localized, Composi- tional Video Question Answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering TVQA: Localized, Composi- tional Video Question Answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.924883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.552668Z digest=sha256:45444b24e50b8c67a46f5a6a8a249c2a31dd0e109c1a1b0d6592b8084669b3ac

Observation 5923f293-cfe4-4b89-a1ea-0378f6ec0dc3 · outbound

This paper cites Learning visual question answering by bootstrapping hard attention.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning visual question answering by bootstrapping hard attention

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.890519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.560966Z digest=sha256:1fc9ddb43033f9e816c9d930d8795b743bc26ab4aaccfb3276d57eb1ce5563cf

Observation 4662b04d-38d0-4ce2-a46a-e98bae3fb0ee · outbound

This paper cites Benchmarking Classic and Learned Navigation in Complex 3D Environments.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Benchmarking Classic and Learned Navigation in Complex 3D Environments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.568564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.568564Z digest=sha256:d6af2c888164831f61e3bca668c6ef0bbc6af19e16ce465ba1452a55514b63a1

Observation dc25097c-b4e1-4eab-85ef-b28af294b1f3 · outbound

This paper cites MarioQA: An- swering Questions by Watching Gameplay Videos.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering MarioQA: An- swering Questions by Watching Gameplay Videos

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.867007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.583764Z digest=sha256:ffc5e63114a083f0cb32d6ea6b8ba21779ee372f30732e3e7fb2f4a65775893b

Observation 07c20715-2e7c-4fac-870e-9046ed3ed624 · outbound

This paper cites Out of the box: Reasoning with graph convolution nets for factual visual question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Out of the box: Reasoning with graph convolution nets for factual visual question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.825665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.596086Z digest=sha256:eb33fb2d102a93489e2aafc9d542097f3154736a9e5ce5b51a7a16c2bc4422dc

Observation 96e187f4-df2d-44e5-a2a6-c47974ab5ccb · outbound

This paper cites From FiLM to Video: Multi-turn Question Answering with Multi-modal Context.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering From FiLM to Video: Multi-turn Question Answering with Multi-modal Context

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-14T13:32:25.943888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.601770Z digest=sha256:d6e7d83627e266aaa75c5bbf6bdcf66cca7c96b9e06a2639a7b8db839d76db54

Observation 7774f0a4-0872-4422-a644-259a5252035c · outbound

This paper cites Learning conditioned graph structures for interpretable visual question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning conditioned graph structures for interpretable visual question answering

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.789387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.609549Z digest=sha256:c09f0fe34d4bdbfa9a867a293efb1ae89d9b23017a3f75c67bc25d4db36871c4

Observation fd108471-39a0-4854-aa2d-fd978956a179 · outbound

This paper cites FiLM: Visual reasoning with a general conditioning layer.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering FiLM: Visual reasoning with a general conditioning layer

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.759038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.616696Z digest=sha256:256d6410067a0f7b52e7aa2824b495d085f254255a840e1b3d2f9eb016a96a87

Observation 9b2716db-af14-4eb8-b7d3-59523db26a58 · outbound

This paper cites Exploring models and data for im- age question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Exploring models and data for im- age question answering

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.730829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.625152Z digest=sha256:9c59d718fc5c455b3d92f6313f506b621a49856194f34377a6873e5e84777693

Observation 4034e400-0c87-4e4e-ab0b-4fb1589ae54e · outbound

This paper cites Faster R-CNN: Towards real- time object detection with region proposal networks.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Faster R-CNN: Towards real- time object detection with region proposal networks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.698989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.632208Z digest=sha256:462743c3792f7d3fb361ab7f27ab4d63729b1e369142ff3fc185c5406f4b2290

Observation d79b820c-bf94-4b62-b008-a4ae3d148811 · outbound

This paper cites Habitat: A Platform for Embodied AI Research.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Habitat: A Platform for Embodied AI Research

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.640466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.640466Z digest=sha256:899143247762d43b0265eb557e099076924aaf1912e0a7ce43ecbfbfee4c38f9

Observation 609d4e1c-5610-49f8-b2b0-dd44ac332433 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.650423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.650423Z digest=sha256:e2965e6d08653208a771da31eefdba0bf58627fd14b44f38e5ceeec4a40e15b3

Observation 45dde7b6-6029-4324-8100-1548eac231b3 · outbound

This paper cites Semantic Scene Completion from a Single Depth Image.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Semantic Scene Completion from a Single Depth Image

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.651023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.658370Z digest=sha256:ac959de946215f81a7edcfc50704f8c2316f61316ebf3892a96bf6c2d4f53f78

Observation 9c4266da-d693-4479-aa03-dfee70061caf · outbound

This paper cites Visual Reasoning with Multi-hop Feature Modulation.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Visual Reasoning with Multi-hop Feature Modulation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.607478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.672201Z digest=sha256:6868b6bf3ee285db7577c79b09f79845fd131607d47e19f240b7b1b25329ae6f

Observation 968ea279-622b-4d55-8611-1bfe24ce2d6c · outbound

This paper cites MovieQA: Understanding Stories in Movies through Question- Answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering MovieQA: Understanding Stories in Movies through Question- Answering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.579054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.681109Z digest=sha256:1a69792dfd60f52a98c6502495109475d9757e1005749f03b6fd13eaa8f496c3

Observation 77497a4e-c0b6-4329-af56-0812a0ce3c07 · outbound

This paper cites Graph-structured repre- sentations for visual question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Graph-structured repre- sentations for visual question answering

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.546012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.691890Z digest=sha256:4048fbaa5e6d4e80653d2e66669104dae4cf581b8fb3a785d80b6c17fc919f79

Observation c9a2c48c-af4e-4097-ba01-88cb44e7b29d · outbound

This paper cites Learning spatiotemporal features with 3D convolutional networks.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning spatiotemporal features with 3D convolutional networks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.515729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.709462Z digest=sha256:648cc790ed9d7846d7f069edd578c5b8496e152553247148c932aebae44e3202

Observation e0f33697-411a-4f17-815f-210133603017 · outbound

This paper cites Embodied Question Answering in Photorealistic Environments with Point Cloud Perception.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Embodied Question Answering in Photorealistic Environments with Point Cloud Perception

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.483075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.716852Z digest=sha256:ae1b834bad1eff7cba84cb1fd772b8ca68b16a8f34beb15f703c4bb1eb21370d

Observation ee8e3cb8-e2a5-4a54-9dec-db60094833bc · outbound

This paper cites Building Generalizable Agents with a Realistic and Rich 3D Environment.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Building Generalizable Agents with a Realistic and Rich 3D Environment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.726025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.726025Z digest=sha256:375a3838ebfce66cefe12e83fc053878109825ce2668f655ccfe81a60cdce23d

Observation b8195918-9c8b-482b-a52c-36c9ccd4362f · outbound

This paper cites Stacked attention networks for image question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Stacked attention networks for image question answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.733408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.733408Z digest=sha256:c2b144a6bab07634622ed3bd8dce2a61a044d49111ac45b091068ab7993e804f

Observation 416e6c67-7bb7-461b-8dc7-a15cb969a60d · outbound

This paper cites Berg, and Dhruv Batra.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Berg, and Dhruv Batra

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.425804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:32:25.742010Z digest=sha256:cd7e6586c8d3a47262359cc1d92c467d57d977aef3187ebd2f6baf36255c9db6

Pith citing papers

No inbound Pith citation observations are available.