Pith. sign in

Paper Citation Record · LEDGER

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation

As of 19 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 1 inbound Pith citation observation for arXiv:2504.16516.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16516 v2

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:06:08.650648Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T15:39:59.878169Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

72 of 72 outbound references displayed

  • verified exact16
  • verified fuzzy0
  • unresolved51
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ad2bde14-9167-4cfa-860e-2bf1f689dd7a · outbound

This paper cites an unresolved cited work.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:06:13.402502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:07.466374Z digest=sha256:e014ccd1e267b0044513fe41ee2af3c2712eb08fb8efe91e9ad252601ac7eb03

Observation d229a878-0958-4f5e-ba05-64aa8e326b0c · outbound

This paper cites BEVBert: Multimodal Map Pre-training for Language-guided Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation BEVBert: Multimodal Map Pre-training for Language-guided Navigation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.499847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.499847Z digest=sha256:0c5caf3b16ae5e115cd6e69074830ee3c6803f3f55362bbef1832f0f22449536

Observation cbbb6cf6-be96-4af8-890c-d25497dd732d · outbound

This paper cites ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.519988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.519988Z digest=sha256:eda872f339fb1d4f41cc1667e25f5a2d9922a5787aae6fe2004f702950eee77e

Observation 1699d242-f2c4-4073-bb44-5310da61a438 · outbound

This paper cites Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.545272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.545272Z digest=sha256:d48ef7b3410ac75c7f6c9f63048b4b6cd01d385ecaa7eb54ce2f5120f26f4b29

Observation 83d84284-080b-4280-ac67-b3196a929cf1 · outbound

This paper cites an unresolved cited work.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:06:13.373802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:07.567473Z digest=sha256:24027ab5232fcbc62073367ea71c1c8ab177f8a1f111ec010db70cf885a871fc

Observation 76b781b4-2837-441b-8948-d3e778045a1d · outbound

This paper cites Neural Topological SLAM for Visual Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Neural Topological SLAM for Visual Navigation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.580851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.580851Z digest=sha256:c70e4bf4800b32d24fd7bca4d56c2b083b4a8496368114d4bcd60de991e51047

Observation 07456b5e-b80a-434b-b619-dce90441f859 · outbound

This paper cites $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation $A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.603723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.603723Z digest=sha256:2de0ea00b85357e944994ec89014e4c4da8a39560b6e342b9c19e411310ead1c

Observation b4b2cf3f-f1bd-4a7c-883d-2e17e46ea3b5 · outbound

This paper cites History Aware Multimodal Transformer for Vision-and-Language Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation History Aware Multimodal Transformer for Vision-and-Language Navigation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.655484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.655484Z digest=sha256:5bdcc44aebc69e5e6b601ec4df49189d7ce7a3f712b8cde75c18d2aaf3173713

Observation b67a72b3-0421-4514-baf4-26aba5c75063 · outbound

This paper cites Learning from Unlabeled 3D Environments for Vision-and-Language Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Learning from Unlabeled 3D Environments for Vision-and-Language Navigation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.689040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.689040Z digest=sha256:2d1887b537ad5e92e2807ba222d3c67741bcd97c5859654f913d47d2e3790f4d

Observation b8b2441b-e7b6-41ef-8a64-1ddf14674b56 · outbound

This paper cites Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.725653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.725653Z digest=sha256:9bb57587d0b090272611ff017f7505f33c29131dd2a90b065560f0baea31abef

Observation 1e1ee716-811f-4b31-9131-ffe850d2293c · outbound

This paper cites Learning a Recurrent Visual Representation for Image Caption Generation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Learning a Recurrent Visual Representation for Image Caption Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.735728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.735728Z digest=sha256:8cba039be28c8f79f92e824e9c50ad979e39e1cc3bd094dadb49c1c18d008793

Observation b613b955-a92a-412f-9dbe-a47d9e473c88 · outbound

This paper cites an unresolved cited work.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Unresolved cited work

Reference 12

Resolution
verified exact
doi, observed 2026-08-16T11:06:08.855981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:07.762304Z digest=sha256:ea3cdc15127a994d215c9349904a39b6610303e0ba9a92af4de4792e7ac3e474

Observation f3e73971-4aad-47a2-bab0-f63a60d3e4dd · outbound

This paper cites an unresolved cited work.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Unresolved cited work

Reference 13

Resolution
malformed identifier
doi_truncated, observed 2026-08-16T11:06:08.821752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:07.773696Z digest=sha256:4ed79338f85c1ee12fb956bf614015314563a8d9feae631dac03170923778d9a

Observation 1d672bb8-3e73-4dad-bdb0-6fb7610b9bbd · outbound

This paper cites Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.781958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.781958Z digest=sha256:f42c9981fd988b2f9e5cc4e72cb90a9ba0db31c591d4e81f182ee3f1d46dfadf

Observation 1d5d2ba6-6fd8-4009-bbf8-223048b0fe3d · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.795967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.795967Z digest=sha256:a37970f8c421ea15f44438cdb5e120368c79d0b47e17fbfac048504341307505

Observation f785bb5c-c4c0-4b72-a556-09b5d95092fa · outbound

This paper cites Attention as an RNN.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Attention as an RNN

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.807285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.807285Z digest=sha256:8b1412634873ca72d1093dda21456a1464a9a7b2922d418b5317221ede9a090d

Observation 1f074f1b-cb63-4c14-a2f7-a12f78cb470d · outbound

This paper cites Speaker-Follower Models for Vision-and-Language Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Speaker-Follower Models for Vision-and-Language Navigation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.825790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.825790Z digest=sha256:d39227041bfbb7264d244f9641c7e85b2b7cf302a074ff18addfe61c30e99f0c

Observation aa50d21b-507f-4a52-9e4a-7d3339078550 · outbound

This paper cites an unresolved cited work.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:06:13.328395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:07.840213Z digest=sha256:e5d3ebe016938e804da80273db8b4057e1dadcb1d73ed67c3cdbe39fce34f614

Observation 261cb7a2-f8e1-4b8a-8998-531c649ba286 · outbound

This paper cites Airbert: In-domain Pretraining for Vision-and-Language Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Airbert: In-domain Pretraining for Vision-and-Language Navigation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.854228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.854228Z digest=sha256:8e59abc46898fb13dceb5d0d614ced6b8c72ea5505ed590111d02e09f29d8986

Observation 21de8b16-d1cd-4f1c-bdae-572ce1378f5d · outbound

This paper cites Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-training.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-training

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:06:11.921926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:07.868524Z digest=sha256:dcd97f426176fc622c76f6e350933cbdfa1d06bd565954a9776c0fd1a662dfa7

Observation 135980a7-2e73-48fe-97d7-4637a5354f4c · outbound

This paper cites an unresolved cited work.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.879964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.879964Z digest=sha256:de3f8d2d1cc03750ec5ffe1399e35178d1b6ae231f0deb91b0f99707f6354675

Observation 94a5d31c-75a6-4065-a4b7-39200e33d59f · outbound

This paper cites A Recurrent Vision-and-Language BERT for Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation A Recurrent Vision-and-Language BERT for Navigation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.911540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.911540Z digest=sha256:4a5cede12be4266078af2d20257dd1be8611f9e5b7a5c915b5d2b16a39ac136d

Observation b8b33199-7c83-4085-ad74-7e74a9f2280f · outbound

This paper cites an unresolved cited work.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Unresolved cited work

Reference 23

Resolution
verified exact
doi, observed 2026-08-16T11:06:08.775389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:07.926090Z digest=sha256:88fe7e748200fe8a6813c65031d0a9f66b72f4a9cd4599c98a03c9ce50ac03b2

Observation 4783f1ff-e4a5-481b-82a8-8e03baaf892e · outbound

This paper cites an unresolved cited work.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.950393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.950393Z digest=sha256:32ef9ade888e3c03b8e8aca1104f048175c38b42f6b9aea72867640061e28375

Observation e50d9ea0-9b8b-4ff8-b1ff-33146436917b · outbound

This paper cites Path Planning and Navigation Inside Off-World Lava Tubes and Caves.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Path Planning and Navigation Inside Off-World Lava Tubes and Caves

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:06:11.695926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:07.976483Z digest=sha256:fc26e1f08c80354aa9af2dd70948e777671b2439b8d4ba5f46aebd0851870d48

Observation 0d7707cf-d0f5-4175-a900-d79c67ba6587 · outbound

This paper cites Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.991028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.991028Z digest=sha256:de839a2c79ff51a99520b2d5be0470d7b7c833329e7de6bcd2a3b7b49a0ad039

Observation 70f74848-cae4-4864-ad8c-109592e8c126 · outbound

This paper cites an unresolved cited work.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.006053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.006053Z digest=sha256:37fd2049f6355d06a5ee2d226bddd6afb0561d636965ef635adc446d60f7b76b

Observation 441333e6-66fd-4831-b682-33db62a7fdc3 · outbound

This paper cites KERM: Knowledge Enhanced Reasoning for Vision-and-Language Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation KERM: Knowledge Enhanced Reasoning for Vision-and-Language Navigation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:06:11.471905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:08.031009Z digest=sha256:532e1675b0ac0cf48994d91e328911b2796957677ad5bef5f308c79339beba20

Observation bcc5dd8e-d0d9-4137-b4b1-7017aaa5eea9 · outbound

This paper cites NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.046318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.046318Z digest=sha256:23666c1d8c59dc05b69c79e8919e2246d122e2b05199557e42a536e640018482

Observation 84fb2f0d-4f24-4747-8d8b-23609b7472e2 · outbound

This paper cites Recurrent Image Captioner: Describing Images with Spatial-Invariant Transformation and Attention Filtering.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Recurrent Image Captioner: Describing Images with Spatial-Invariant Transformation and Attention Filtering

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:06:11.285536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:08.068118Z digest=sha256:4a032b9353674963c6f4fa335fbb2b86e20cd523ef56a9025f03a151dd16ca8c

Observation 57c3f1ec-02a5-414d-bbcf-e26b1a9c6a87 · outbound

This paper cites A Unified Virtual Mixture-of-Experts Framework:Enhanced Inference and Hallucination Mitigation in Single-Model System.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation A Unified Virtual Mixture-of-Experts Framework:Enhanced Inference and Hallucination Mitigation in Single-Model System

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:06:11.216778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:08.075253Z digest=sha256:4606d36e669f5bbb5c46060a56d8b3386c2dbf677f7a74334118a8055e100b20

Observation b2c26202-c90a-4a09-ab8f-8ae360e15ff8 · outbound

This paper cites Discuss Before Moving: Visual Language Navigation via Multi-expert Discussions.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Discuss Before Moving: Visual Language Navigation via Multi-expert Discussions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.081330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.081330Z digest=sha256:f533cc65de70bb77f26f756b1107d7d0aafb7b2a7ce00af1e619a3420b2df65d

Observation f2c6313f-9df0-47cc-8424-173814867035 · outbound

This paper cites an unresolved cited work.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:06:13.191179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:08.095462Z digest=sha256:519bdfdd65b514ec99de932fe899feeb72ac63124acf27329825937e04500664

Observation fd72d41c-dd0e-494e-adfb-c45570de4d55 · outbound

This paper cites Object-and-Action Aware Model for Visual Language Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Object-and-Action Aware Model for Visual Language Navigation

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:06:10.970570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:08.109456Z digest=sha256:c239e233d157997a4b11472d44fd00d4b0c3c05c0f4731beaf21a5d7429c0374

Observation 831aa656-4a58-4d4a-afb6-dd09264d5213 · outbound

This paper cites REVERIE: Remote Embodied Visual Referring Expression in Real Indoor Environments.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation REVERIE: Remote Embodied Visual Referring Expression in Real Indoor Environments

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.116380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.116380Z digest=sha256:0ef09124b63b43ca41984db7c728496441d0b004d65a49ebadbfc9f0983682b6

Observation 3b2f1139-704a-4bd0-8d94-b0e4ca4b7d45 · outbound

This paper cites Recasting Generic Pretrained Vision Transformers As Object-Centric Scene Encoders For Manipulation Policies.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Recasting Generic Pretrained Vision Transformers As Object-Centric Scene Encoders For Manipulation Policies

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.123049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.123049Z digest=sha256:364cfb5d80f66be1ebeec82440b6df011f5b6fddae9214b7686a774416600e57

Observation c9e23dcf-6252-45a1-9cc1-069232b0e70c · outbound

This paper cites SOAT: A Scene- and Object-Aware Transformer for Vision-and-Language Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation SOAT: A Scene- and Object-Aware Transformer for Vision-and-Language Navigation

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T11:06:11.043438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:08.103749Z digest=sha256:f307a03195a94f7c6094e6c4a0a1f8baad40c351f858b8ed145d1de19f149a3f

Observation 06e1335e-548d-41ef-8465-63e11fc42636 · outbound

This paper cites March in Chat: Interactive Prompting for Remote Embodied Referring Expression.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation March in Chat: Interactive Prompting for Remote Embodied Referring Expression

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:06:10.710812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:08.167772Z digest=sha256:406bb378bfc7370bad03ca8900b12cdd686e8b0d6f67c8d270ecdb9fbb2772dd

Observation f1a0d447-2f78-454d-ac06-a1c6046ac1e3 · outbound

This paper cites VLN-PETL: Parameter-Efficient Transfer Learning for Vision-and-Language Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation VLN-PETL: Parameter-Efficient Transfer Learning for Vision-and-Language Navigation

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:06:10.623312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:08.178251Z digest=sha256:c28f420095c6bf807b0f8ed2f912370a20aebf2e67b98ce93ff24d0a6d9e2882

Observation 0bb0f6c5-5ca4-4b97-a7b3-37e6fd35d0e6 · outbound

This paper cites an unresolved cited work.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.188880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.188880Z digest=sha256:c9a1c704aabce537c9f73950dce261304f61ab821da7f711f23806cce17f9346

Observation ecdcfc8d-187e-4d55-a105-b675384dcf3e · outbound

This paper cites an unresolved cited work.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.134321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.134321Z digest=sha256:c176f3d13ab2bf27450a8c53cd726227c60cf8d3d7eb1d55a72fe342c110c6b5

Observation e9681660-397b-4e88-bc12-9d36904536f4 · outbound

This paper cites HOP: History-and-Order Aware Pre-training for Vision-and-Language Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation HOP: History-and-Order Aware Pre-training for Vision-and-Language Navigation

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T11:06:10.766474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:08.146056Z digest=sha256:943416d5542241764d8182aba4457dd9d2a82032c5255548c3a8571481d0a93e

Observation bf5812d1-e8c3-4b4a-9b7f-0d30a7e85e4a · outbound

This paper cites an unresolved cited work.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.219946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.219946Z digest=sha256:b738f4d8f597ba64336443989f0f6c80c25a434a885d5d586a1d8c2eb1c21dae

Observation 97c0c064-284a-460a-8817-3ea8b02238a3 · outbound

This paper cites How Much Can CLIP Benefit Vision-and-Language Tasks?.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation How Much Can CLIP Benefit Vision-and-Language Tasks?

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.225352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.225352Z digest=sha256:892b3a796cc8eda9034d0aa5938058d2e5cffec0d4b6fba6c6134957ca05c889

Observation 37a1bb22-d106-4120-a957-9c4eda0dc0a0 · outbound

This paper cites Noise-Resilient Symbolic Regression with Dynamic Gating Reinforcement Learning.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Noise-Resilient Symbolic Regression with Dynamic Gating Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.233759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.233759Z digest=sha256:b730a0533bd6b1642d83f6e42883456d7fb8546fe617c562d614c162d5285984

Observation 7f6650dc-ae82-476c-bc12-dc36c3f99b6f · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Learning Transferable Visual Models From Natural Language Supervision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.198688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.198688Z digest=sha256:86472fcd9659f0e081024dbed71c907d6cbee7637d54c02ecc565d8696d7ea66

Observation 357441f2-3c70-4317-a7c1-c3da450c1bb3 · outbound

This paper cites Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.214231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.214231Z digest=sha256:6838e51ef5e786be030e9bc6bd8551824f3621b3ff6790db2308abfe47ee2abd

Observation 0e56d0db-52be-481a-9240-371285628e17 · outbound

This paper cites LXMERT: Learning Cross-Modality Encoder Representations from Transformers.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.271871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.271871Z digest=sha256:aaba90f923f1b502f161ff4aa270acbfa9274d6dc43322dc74e16e8e5cbb24f3

Observation affd29ba-7875-4ed6-91e9-7dc93c5db6cb · outbound

This paper cites Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.287416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.287416Z digest=sha256:a4945c1a1381ecc1819198233c52c55ae697fcddef95b2475514d24be0d6523e

Observation 11977878-68ad-45d8-b322-28adfcc7d736 · outbound

This paper cites Structured Scene Memory for Vision-Language Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Structured Scene Memory for Vision-Language Navigation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:06:09.830680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:08.303001Z digest=sha256:42bc455f282b0b078425384aa5452db0de1616737b381c63335dfbf16ca27b45

Observation 96bb6cb5-1150-4a8f-a9c5-dbdff674ac49 · outbound

This paper cites CRUR: Coupled-Recurrent Unit for Unification, Conceptualization and Context Capture for Language Representation -- A Generalization of Bi Directional LSTM.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation CRUR: Coupled-Recurrent Unit for Unification, Conceptualization and Context Capture for Language Representation -- A Generalization of Bi Directional LSTM

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:06:10.049623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:08.239929Z digest=sha256:36039ee0b3ce671c65ca43f71d032d6ddc9794c867229c7b5a1a954d859af9ba

Observation 6ed0e2ee-181d-415e-8439-11c324d558ea · outbound

This paper cites Scaling Data Generation in Vision-and-Language Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Scaling Data Generation in Vision-and-Language Navigation

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:06:09.679691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:08.353251Z digest=sha256:3bb5eb41832abed0e523f777ee2dc10461c5a1166e9489f22dff9d3a8194d23c

Observation 72cdfa28-fe1a-407c-a2db-9842fb1329ea · outbound

This paper cites an unresolved cited work.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:06:13.033549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:08.374966Z digest=sha256:c623725e67054076e9ace14798a5df01a040b13cd0245b084835cc0d9f3a4a0e

Observation 4c801cb5-4993-4f15-936b-2da4a8541fb8 · outbound

This paper cites SimMIM: A Simple Framework for Masked Image Modeling.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation SimMIM: A Simple Framework for Masked Image Modeling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.394168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.394168Z digest=sha256:c7bc3a0b1dd82dbc41380fcbfb154e003ddcf7e84bd53eceee088bfe4156b973

Observation 4cd00b6b-5246-4199-8bc2-fbcce4a5c538 · outbound

This paper cites Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.308245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.308245Z digest=sha256:3d4b43e83b4bffdf78e2ba6b7a9a740582d50ab8939473ba8ff69fc6b2b1f20a

Observation a1be83a5-e4b6-423b-b63f-e15f0fe11d79 · outbound

This paper cites Lana: A Language-Capable Navigator for Instruction Following and Generation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Lana: A Language-Capable Navigator for Instruction Following and Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.331220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.331220Z digest=sha256:8d977c8745bd2797a7d8b52beeb7772645e371f206cf2c0e34f51f707cb4b306

Observation 1510ace9-ad89-4eae-a863-3734039aa2d6 · outbound

This paper cites CREStE: Scalable Mapless Navigation with Internet Scale Priors and Counterfactual Guidance.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation CREStE: Scalable Mapless Navigation with Internet Scale Priors and Counterfactual Guidance

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.441977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.441977Z digest=sha256:d959ec08638cd65e274b7fcfff501486d49a1f4ab253b45a52c65f961f78d930

Observation 62eb3bd5-ec20-4cf0-84e3-42031a9cc4a4 · outbound

This paper cites Mind the Gap: Improving Success Rate of Vision-and-Language Navigation by Revisiting Oracle Success Routes.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Mind the Gap: Improving Success Rate of Vision-and-Language Navigation by Revisiting Oracle Success Routes

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:06:09.381039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:08.472362Z digest=sha256:b7890235a24d1ba645573949ae8b631224f9ecd0832f9953804a8d0bd50ec5cc

Observation 16abc862-54ea-4c51-a9d6-2597265ca0c4 · outbound

This paper cites Towards Learning a Generalist Model for Embodied Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Towards Learning a Generalist Model for Embodied Navigation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.490229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.490229Z digest=sha256:7eb58d89f899816a73789191b92360bf07bf80d0a15dab384f537ff3e62fa684

Observation 680dd3e3-6a97-4acd-81c5-c69ce9912374 · outbound

This paper cites NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.513280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.513280Z digest=sha256:cf15f9ffccf69cf801cd8c5c64568b7318fccab468b88073038856d041c7ad92

Observation c78a99d6-3d91-4e8b-8325-1d2a8282ff18 · outbound

This paper cites HorNet: A Hierarchical Offshoot Recurrent Network for Improving Person Re-ID via Image Captioning.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation HorNet: A Hierarchical Offshoot Recurrent Network for Improving Person Re-ID via Image Captioning

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:06:09.531885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:08.403390Z digest=sha256:68f6509a450100c7fdcf9deb51786036d325a76d05df3c6b344670b86382f5b8

Observation 629a6ed9-0ff3-4da5-8715-39635f527555 · outbound

This paper cites an unresolved cited work.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.434649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.434649Z digest=sha256:f60dcb4f2c7940ad51946c65c489981fa9999fbe7c6285954f62629d1035bfbf

Observation 65e0daa8-089d-490b-b588-03f8bd86e74f · outbound

This paper cites Vision-Language Navigation with Self-Supervised Auxiliary Reasoning Tasks.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Vision-Language Navigation with Self-Supervised Auxiliary Reasoning Tasks

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.553062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.553062Z digest=sha256:9df5e0d840de563aa1a613bb8bf2dccd94c8656337f9f585d100e336336ec873

Observation 921ead16-f975-43fa-803e-44b0590a008e · outbound

This paper cites MiniVLN: Efficient Vision-and-Language Navigation by Progressive Knowledge Distillation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation MiniVLN: Efficient Vision-and-Language Navigation by Progressive Knowledge Distillation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.564391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.564391Z digest=sha256:3e7322be11deb542e38be8736e0b8bebc3beabbc27156a852c6ea354bccc808a

Observation 9c2a6e2b-8c09-48ca-8185-22789570984a · outbound

This paper cites an unresolved cited work.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:06:12.933474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:08.618907Z digest=sha256:0d1022b1886ce459b1dfc3026f1beed25e12b8aaef1132d41c03d565dd7c4145

Observation 04ac7572-bd40-4b9a-937b-006744c2040e · outbound

This paper cites Rethinking the Spatial Route Prior in Vision-and-Language Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Rethinking the Spatial Route Prior in Vision-and-Language Navigation

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:06:09.134252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:08.523214Z digest=sha256:f7c3d51c602edf44407525d1a00a3631d52e2f165f303e2ebf23b35b4f213437

Observation e9637457-7f42-4258-aab7-3e97b9120167 · outbound

This paper cites SOON: Scenario Oriented Object Navigation with Graph-based Exploration.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation SOON: Scenario Oriented Object Navigation with Graph-based Exploration

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.536657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.536657Z digest=sha256:ede6929f022fc6e1bf282e475550b22d4e6584ce88ce23a006297126a4e4db0f

Observation a0266dda-aa5c-4670-9b97-9556c5ee0a81 · outbound

This paper cites Diagnosing Vision-and-Language Navigation: What Really Matters.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Diagnosing Vision-and-Language Navigation: What Really Matters

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.650648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.650648Z digest=sha256:007181433369d32ad6470dbdfa025c35c4d765f7b1306cfc05aa55255b993b48

Observation 1045577f-c177-49b2-9635-45312d44d9f9 · outbound

This paper cites Image Captioning and Visual Question Answering Based on Attributes and External Knowledge.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Image Captioning and Visual Question Answering Based on Attributes and External Knowledge

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.382404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.382404Z digest=sha256:93ad6aa928e448732c85ff76e852e537dd7915c125ffc5341a56ba36deb5c624

Observation e78ba485-dc43-4352-ba34-c8343d972f81 · outbound

This paper cites Language and Visual Entity Relationship Graph for Agent Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Language and Visual Entity Relationship Graph for Agent Navigation

Reference 2020

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T11:06:11.845052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:07.895918Z digest=sha256:209e5984e9d08a942a1a3c5964db85feec6f78ed12070f6be162fabebcde1c60

Observation 1e71d1dd-ec98-4a18-8362-e41e2cf91132 · outbound

This paper cites Neighbor-view Enhanced Model for Vision and Language Navigation.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation Neighbor-view Enhanced Model for Vision and Language Navigation

Reference 2021

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T11:06:12.850825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:06:07.489788Z digest=sha256:429ab7297780e49d108fb6251c1bd1f277f03c50e64181002402ded8e0aef428

Observation f7419997-a4fa-4fcb-a4aa-0dab656b6e5b · outbound

This paper cites In European Conference on Computer Vision.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation In European Conference on Computer Vision

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:07.963783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:07.963783Z digest=sha256:937c405b377e7d4c330d4bf0c5da679058fb8c42be7752d7e58e98503309830a

Pith citing papers

Observation 225ef4b9-dcb8-461f-aef4-fa058b0ea40c · inbound

A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation cites this paper.

A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation

Reference 178

Resolution
unresolved
no resolver link, observed 2026-07-14T15:39:59.878169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:39:59.878169Z digest=sha256:2fd4ac1ade2643fe3dee3c5663ce258c4931e979cc50992bbebd2b40c4548fb3