Pith. sign in

Paper Citation Record · LEDGER

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

As of 20 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 20 inbound Pith citation observations for arXiv:2411.08579.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.08579 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:36:19.687871Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:55:39.925157Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:19:30.523454Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact1
  • verified fuzzy43
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 04fe7347-d2af-4c67-8b3f-6876f353ca2e · outbound

This paper cites Aerialvln: Vision-and-language navigation for uavs,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Aerialvln: Vision-and-language navigation for uavs,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.288573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.524424Z digest=sha256:cc2911fb4f42445985c0842de7e1edb3a7b48e11e2e2f4c01460d11e7817e5cd

Observation d02febe6-863d-4560-b23a-d4476a0d7b9c · outbound

This paper cites Vision- and-language navigation: Interpreting visually-grounded navigation in- structions in real environments,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Vision- and-language navigation: Interpreting visually-grounded navigation in- structions in real environments,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.280133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.527708Z digest=sha256:75918451615f8e2d20490632a69d78e3dcea1dc4089ec0bc227954f7fc389ba9

Observation 50c515e4-f099-49ca-88c6-f8938fb987ec · outbound

This paper cites Reverie: Remote embodied visual referring expression in real indoor environ- ments,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Reverie: Remote embodied visual referring expression in real indoor environ- ments,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.271809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.530423Z digest=sha256:89c00323328455e53a6f02d69dce23e46e2e270042f8d725dbb812ef2626d821

Observation c028d0b6-99e3-438b-9115-b29d8be19baf · outbound

This paper cites Stay on the path: Instruction fidelity in vision-and-language navigation,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Stay on the path: Instruction fidelity in vision-and-language navigation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.263047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.533221Z digest=sha256:a76ab0e3e72e8ff1475609a8c2b8da302342d81bb86b5bb5f81be8dfb5f840b0

Observation 8167c85e-b03d-4658-8d39-7bb3a5ee43ef · outbound

This paper cites Room-across- room: Multilingual vision-and-language navigation with dense spa- tiotemporal grounding,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Room-across- room: Multilingual vision-and-language navigation with dense spa- tiotemporal grounding,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.253757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.536142Z digest=sha256:3bfda91e9f6655f16a80833ff2534e31e25ab63d9fdbfd32607f42b4d69e6a01

Observation 481fcf3f-a758-4e06-946d-c1609a40be6d · outbound

This paper cites Beyond the nav-graph: Vision-and-language navigation in continuous environments,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Beyond the nav-graph: Vision-and-language navigation in continuous environments,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.246465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.538936Z digest=sha256:91f10c55940f2cbd0b726bb96226e2dbc79ea98792ff15a05d5bd6f48d87cf38

Observation 5ff2afa4-1dda-4515-b01d-6d3af43202b5 · outbound

This paper cites Touchdown: Natural language navigation and spatial reasoning in visual street environments,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Touchdown: Natural language navigation and spatial reasoning in visual street environments,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.239046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.541736Z digest=sha256:9b0776c4a266a5729e9fa91bb2d813707867d4b7f1206003f7d7932c4ecd87a0

Observation 6b3e3a5c-0025-4063-8dcb-9ee6423bebfa · outbound

This paper cites Dji drone solutions for inspection and infrastructure construction in the oil and gas industry,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Dji drone solutions for inspection and infrastructure construction in the oil and gas industry,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.229332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.544208Z digest=sha256:3e5f891f3dbf5f8e9ccf93a294fe7e8370bdcd7e1ac174f144956839920429ea

Observation 8558fc49-85bb-474c-b317-86724bd9dc96 · outbound

This paper cites Dji drone solutions for optimizing operations in the public safety industry,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Dji drone solutions for optimizing operations in the public safety industry,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.220369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.546761Z digest=sha256:9ff2491b6ebfa13cdf80a030989084498e6ba9336ed9550dce2d6b391a915660

Observation 55091f4a-2bd7-44c1-964f-5fe7f245fcef · outbound

This paper cites Dji drone solutions for surveying, urban planning, aec, and natural resource management,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Dji drone solutions for surveying, urban planning, aec, and natural resource management,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.211167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.549505Z digest=sha256:bccc68c934b36e019af243502215498ea3a853d9901ac2c8f90529e130805700

Observation 9bc98473-640a-4dbb-99d8-4343722935b2 · outbound

This paper cites Hop+: History- enhanced and order-aware pre-training for vision-and-language naviga- tion,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Hop+: History- enhanced and order-aware pre-training for vision-and-language naviga- tion,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.202399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.551791Z digest=sha256:02038390df5b13f6fb5eb51bfdb8e0844dd0a93b70088bd090301904373e7aef

Observation 14b69d68-b0b4-43a7-b915-71a6246b0d82 · outbound

This paper cites Correctable landmark discovery via large models for vision-language navigation,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Correctable landmark discovery via large models for vision-language navigation,

Reference 12

Resolution
metadata mismatch
raw_fallback, observed 2026-08-12T21:36:19.880078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.554070Z digest=sha256:820f6f95a274705d091edf106990fdd9161182c2e00d108861915562e07216f8

Observation 05b6b6d9-d36c-44ca-bdc5-ad72ac5e78dc · outbound

This paper cites Learning to follow and generate instructions for language-capable navigation,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Learning to follow and generate instructions for language-capable navigation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.193474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.556607Z digest=sha256:f545bb0592b18feff507e782c38b89d59bd140da2dfbe7a75e0e7ac86656dcc1

Observation ac166ce2-4757-4d7c-b6e3-40c2a72d9b54 · outbound

This paper cites ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.559058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.559058Z digest=sha256:8bc8cb2c6d7b17997eb914a9e70ad46b1210d6a2b86e4a6e9e03c84f9895bc6b

Observation be234d5e-41ea-4056-acee-55b1c92de0d8 · outbound

This paper cites Towards deviation-robust agent navigation via perturbation-aware con- trastive learning,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Towards deviation-robust agent navigation via perturbation-aware con- trastive learning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.183382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.562158Z digest=sha256:d0c6a90e3ba7b8c60e04ec6644404eac848f9359e4ede54d7840e85a694350f8

Observation ed55ea5f-55fd-443b-ad70-5c8664cf7d43 · outbound

This paper cites A survey on vision-based uav navigation,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation A survey on vision-based uav navigation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.173548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.564603Z digest=sha256:933dcfa9637142cd0d9637c11100ea0b491d6e49167f8f7de7e3ab78041d4ea9

Observation 8b95d68b-795b-46f2-977f-e43f83ed39ee · outbound

This paper cites Vision-based navigation of unmanned aerial vehicles,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Vision-based navigation of unmanned aerial vehicles,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.163859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.567166Z digest=sha256:dca4e22bf8be20df5ebbdae344d1016a945ed8f4d3506beac45abd5a741b2c39

Observation 87158f5f-f303-41c3-9aae-244024c4b5cc · outbound

This paper cites Mapping instructions to actions in 3d environments with visual goal prediction,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Mapping instructions to actions in 3d environments with visual goal prediction,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.154605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.569665Z digest=sha256:4064cc7f35db854a1949d7e72da7ef617a09093d220f9ff754cfda2e27805121

Observation 3dfead87-f1b0-423d-b421-8bc819ef90a1 · outbound

This paper cites Following high-level navigation instructions on a simulated quadcopter with imitation learning,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Following high-level navigation instructions on a simulated quadcopter with imitation learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.145386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.572086Z digest=sha256:10d433f6cc36ef904a70ca088de0a414733867b56fb2861edbfbdcc6c3290bf3

Observation 740247c1-d6d0-4488-acc2-50780c0f342b · outbound

This paper cites Mapping naviga- tion instructions to continuous control actions with position-visitation prediction,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Mapping naviga- tion instructions to continuous control actions with position-visitation prediction,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.135601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.574502Z digest=sha256:0131a260a9a01e693d86abd116f70842f9576098c7a79a2476740bd263a92781

Observation 5c176195-5011-4465-afe1-9ca700386462 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.577236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.577236Z digest=sha256:438db5a342bfd6062ed93f473154faadef85b67cdb5f4a55724f2d422b472119

Observation 08c75f42-d292-497b-aa46-88e384b5241a · outbound

This paper cites Visual Instruction Tuning.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Visual Instruction Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.582527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.582527Z digest=sha256:61ddf193e465965bd3b4f0307f52c1e679f7b711465ea091545f0792b8e370ca

Observation 8081e15b-0bed-4e6a-8e0f-a7ef0b1f98f4 · outbound

This paper cites Grounded language-image pre- training,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Grounded language-image pre- training,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.111581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.585270Z digest=sha256:5d40fc884e7377078799443d548121615d86640832e7a26bf9a8ba61f1fe28c7

Observation d5640189-f98a-43fb-ad4d-01639967cd94 · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Instructblip: Towards general-purpose vision- language models with instruction tuning,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.587760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.587760Z digest=sha256:eb4084747f7a1684077724afecabcfacbd303dfcd470875b47d0cbd3f07159b3

Observation b3ff0a49-d507-4da7-bbe7-e2ccbd7c251b · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.590227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.590227Z digest=sha256:7aa5815e9dab17b89fd3170fd554e308dbd6841461fed6cba50ac82d8a2147c4

Observation cd3e0007-e2bd-4113-9be9-ea1546bf5de0 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.592961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.592961Z digest=sha256:e7c5afd9182e8194f892e219bd9bc531fcaa66929c1f24d9497c4e42059a3471

Observation e42cd9e9-2987-4fd5-b6fe-175616d3c02d · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.595607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.595607Z digest=sha256:1dd90148b5c79a017d6a0f302e15835ebc54f7c7024fedae92d6fce59f70ce23

Observation e700f938-6099-49c1-9cc8-9ad87223acd7 · outbound

This paper cites Qwen Technical Report.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Qwen Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.598472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.598472Z digest=sha256:64921dc26f603c1e423c326858bba736a1db10f4322c9435ae300e9d68072e3b

Observation 361d3cf7-3a01-4b60-aaca-0466c615004e · outbound

This paper cites Gsv-cities: Toward appropriate supervised visual place recognition,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Gsv-cities: Toward appropriate supervised visual place recognition,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.097195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.601061Z digest=sha256:b23ee0aaaa2f9932d4e8795674cedc5ac379fc0d94d1a3206b7a072f68415228

Observation 1faea57a-3bfb-4221-9c99-11a2f445dcb8 · outbound

This paper cites The streetlearn environment and dataset,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation The streetlearn environment and dataset,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.088346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.603333Z digest=sha256:6704c2fc857d3f1146df1564d0a7475c57023e92f9b126e26977eedb4b7aa61c

Observation 5c2dcecb-229a-40f0-802b-a76f9f343137 · outbound

This paper cites Learning to follow directions in street view,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Learning to follow directions in street view,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.079866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.605868Z digest=sha256:aa16ea99c8952a3d32a936e7490dd6c35d8ac10bd14e7524914c41018afe881f

Observation a88d4c06-7bf1-4479-ba47-f36f1c8985cd · outbound

This paper cites Retouchdown: Adding touchdown to streetlearn as a shareable resource for language grounding tasks in street view,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Retouchdown: Adding touchdown to streetlearn as a shareable resource for language grounding tasks in street view,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.070836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.608453Z digest=sha256:56646b731347b5ce98a2ccd1bea8680e44fd1e14e3ce8e19b12bcacaec5d2eb8

Observation 27ec9b46-ea08-4b44-912e-985f984d3931 · outbound

This paper cites Silg: The multi-environment symbolic interactive language grounding benchmark,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Silg: The multi-environment symbolic interactive language grounding benchmark,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.061931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.610952Z digest=sha256:d52b1e73c3b6d09b21b2ef06222d758109463e62a8fbdbf05945eb7177676ecb

Observation cb78d08c-45a3-49cd-b72d-c196970336c2 · outbound

This paper cites Outdoor vision-and-language navigation needs object-level alignment,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Outdoor vision-and-language navigation needs object-level alignment,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.053117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.613415Z digest=sha256:0568d0afb418a4f89c6680b4656b37234268c7d4d63d5ccded920f037d1070de

Observation ebeb6a68-5332-4048-95e9-2b7fd5c47faa · outbound

This paper cites A priority map for vision-and- language navigation with trajectory plans and feature-location cues,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation A priority map for vision-and- language navigation with trajectory plans and feature-location cues,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.044234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.615811Z digest=sha256:40236001626cb668e3aa034e50b4ce8b4ceaaaabdf25a35a822a91dc86b2ff8b

Observation c4797cf4-cd03-4f72-8c2d-85c8e909d761 · outbound

This paper cites Multimodal text style transfer for outdoor vision- and-language navigation,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Multimodal text style transfer for outdoor vision- and-language navigation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.035193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.621791Z digest=sha256:c8fa8b392e30389078de8563c53a4afa4919e26b857be05efec34d610fb0d429

Observation 7a35650c-9e21-4afa-884b-6f1c7c6c42bf · outbound

This paper cites LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.624438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.624438Z digest=sha256:6655c74a265e7242538ef16cccf274b772f3b05bf480f3b4120b368c9ab610f0

Observation 9a42f1b5-c8e9-4924-98db-5bf2536bd6b2 · outbound

This paper cites VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.627503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.627503Z digest=sha256:ed9dda1348dafa568b0a7e30787cc3bf112a76068ad656941848a820969b29df

Observation 3f7c5d22-0c64-4ab6-9eea-8ee5d29860b3 · outbound

This paper cites Opt: Open pre-trained transformer language models,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Opt: Open pre-trained transformer language models,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.025980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.630225Z digest=sha256:73c4ceb6c8b4d1b38ee7f18aacd340d4e8852a75af3016f9200bdd6768cc9a7e

Observation b98990a4-4241-4958-a765-dd272f545f0e · outbound

This paper cites Palm-e: An embodied multimodal language model,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Palm-e: An embodied multimodal language model,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.017244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.632744Z digest=sha256:28918ceb89001dfa2beedff9b90e90315b21c98eb2e5c72d701a1c17fbf6d512

Observation 671d8ebc-8939-401d-ae97-9604757d5c70 · outbound

This paper cites Training language models to follow instructions with human feedback,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Training language models to follow instructions with human feedback,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.008623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.635530Z digest=sha256:21d005c689d61ab5f0679759bd2fcae111a4a5ed7888d1d66b425663f11197b1

Observation 68054db1-e704-4b2a-902d-dd2bbbaa29a2 · outbound

This paper cites GPT-4 Technical Report.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation GPT-4 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.638214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.638214Z digest=sha256:bbee7accac2d5a06e57f22a9bf06b0e8c5a12b8d12ad760b708a4db8580bdb92

Observation e30296cb-10a5-4f4f-9065-a72edb665336 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation LLaMA: Open and Efficient Foundation Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.641203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.641203Z digest=sha256:23bdc39d04479726b4ac03d71c61d9568e0b307a02d78197c3963f0b39f22e66

Observation e59f452c-8b74-461c-aecf-fbac03409379 · outbound

This paper cites ESC: Exploration with Soft Commonsense Constraints for Zero-shot Object Navigation.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation ESC: Exploration with Soft Commonsense Constraints for Zero-shot Object Navigation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.643739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.643739Z digest=sha256:f87c9ba9fc18a33e7daea6010dfe5af572c9e0922381409e18e4cdd4c6e31977

Observation 018b6d5a-777a-4d61-b9be-40de66b54966 · outbound

This paper cites NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.646539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.646539Z digest=sha256:1c59d17efa76f5342e9e0d506fc0a6f56d60d78ea31ccefc449ad35cd2c4e064

Observation 981a5aad-1fb3-4393-b193-a96efc1b2a7b · outbound

This paper cites Neural slam: Learning to explore with external memory,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Neural slam: Learning to explore with external memory,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.999617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.649143Z digest=sha256:285968b9f32db953ea99ba792031f0fe111ab1f784b8e5bf32e9c2ab5c88d414

Observation 97d9714a-e0d8-40ed-a4a6-f453f5a2b1c9 · outbound

This paper cites Egomap: Pro- jective mapping and structured egocentric memory for deep rl,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Egomap: Pro- jective mapping and structured egocentric memory for deep rl,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.990743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.651661Z digest=sha256:475ef8177fbba7ec40880af22f4da9ad048ffb2ef09c908c78e9dd0807ff3419

Observation f0dd38f3-aeee-4319-b76c-8bc758bef865 · outbound

This paper cites Semantic mapnet: Building allocentric semanticmaps and representations from egocentric views,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Semantic mapnet: Building allocentric semanticmaps and representations from egocentric views,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.981765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.654249Z digest=sha256:765864d652aea41209c5e0e3c667b3bcefb2ea44d77fee8901e68ebcce833421

Observation 5752096d-59e7-40e4-a8be-53be577cebd1 · outbound

This paper cites Audio visual language maps forrobot navigation,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Audio visual language maps forrobot navigation,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.972719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.656831Z digest=sha256:61674277b27c3daa92d226dda2fcaf5eab94f423320a03e60cb5c62252fdb828

Observation d0e7dad1-cedc-4b88-8ddc-f46816f24fc3 · outbound

This paper cites Mapnet: An allocentric spatial memory for mapping environments,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Mapnet: An allocentric spatial memory for mapping environments,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.963827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.659298Z digest=sha256:caf2b8dc0c2db47816e52e5d47056a1fda57a6265a7702eb245e20b8dc531ccd

Observation 02cfc938-f6e9-4897-98b1-d6ee4feea425 · outbound

This paper cites BEVBert: Multimodal Map Pre-training for Language-guided Navigation.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation BEVBert: Multimodal Map Pre-training for Language-guided Navigation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.661906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.661906Z digest=sha256:1a2284f4193a9c95c7e6afa8266cde5db890d849fcd47df94e8af4793300d471

Observation 9b204acb-cc5b-434c-b783-d05be3d74076 · outbound

This paper cites Cognitive mapping and planning for visual navigation,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Cognitive mapping and planning for visual navigation,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.954209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.664626Z digest=sha256:8c9e214b8153114b7bfb357fdd076b9b21a7b27610a614c52a37823769ac6e39

Observation 031e2c47-b4a9-4626-9b0f-3ffd4591cdbc · outbound

This paper cites Semantic mapnet: Building allocentric semantic maps and representations from egocentric views,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Semantic mapnet: Building allocentric semantic maps and representations from egocentric views,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.943942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.667053Z digest=sha256:f57d04b77538d0496d60961df6c0b59de636508af19002a63918fc5c63db86aa

Observation 041a2fff-21b9-470c-b623-41433d2fff10 · outbound

This paper cites Cross-modal map learning for vision and language navigation,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Cross-modal map learning for vision and language navigation,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.934773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.669525Z digest=sha256:2cec984dc514baa4cba0f15c362d28d7b0dad6cc1f6455992492f43002d2fe2b

Observation 569d789a-e2f9-4554-bfe0-ed6486ca583b · outbound

This paper cites Topo- logical planning with transformers for vision-and-language navigation,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Topo- logical planning with transformers for vision-and-language navigation,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.925391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.671919Z digest=sha256:44a83dd5e4a224a02e460eb4b37fefacbcd0650f20a06f10e6e9c873f031851b

Observation c54bd769-1044-4b50-b5d9-8f4b8bfd62f0 · outbound

This paper cites Generating landmark navigation instruc- tions from maps as a graph-to-text problem,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Generating landmark navigation instruc- tions from maps as a graph-to-text problem,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.915520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.674472Z digest=sha256:852fab60af47230aefb784174fe68456f8fe1148d7188b89d021dd363c695d3d

Observation 81e05be7-04b8-4461-8a59-f34ae2111060 · outbound

This paper cites Microsoft COCO: Common Objects in Context.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Microsoft COCO: Common Objects in Context

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.676906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.676906Z digest=sha256:6bcf9a44ffa7b7757be2353e7973d7eafaef58c7d8b310faf504d6ad1f2003cc

Observation 2e33fcc0-dadf-4a77-ab13-678911f39f40 · outbound

This paper cites Dynamic head: Unifying object detection heads with attentions,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Dynamic head: Unifying object detection heads with attentions,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.906035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.679766Z digest=sha256:23c0c4f6e004197c2ea341cc9d0bfe2615b34107d05f8161f59100985e544e5d

Observation 0b591f49-95c5-4320-b752-ced7d0b05aa3 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.682360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.682360Z digest=sha256:481bcb54d359a3b39ff22ce0476f70a16dd5b70686f9f497ac102d32a1312753

Observation 97605107-7c0b-4f76-95fd-549fe7bb7d30 · outbound

This paper cites Analyzing Generalization of Vision and Language Navigation to Unseen Outdoor Areas.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Analyzing Generalization of Vision and Language Navigation to Unseen Outdoor Areas

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-12T21:36:19.715144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.687871Z digest=sha256:35bc8a0736b2893993206451d11cec4ad88a0e28c045d15478bc8d19af4339df

Observation e20f1ef2-14e0-4d8a-8444-24590a1c84f5 · outbound

This paper cites Available: https: //api.semanticscholar.org/CorpusID: 52967399.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Available: https: //api.semanticscholar.org/CorpusID: 52967399

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.890398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.685176Z digest=sha256:385f4f52b0f6a9bf9dd140bcb9e2d6cbb0cbc6c6a5ef7d23a3ebd17f76de02ef

Observation bcce2f57-31aa-4f75-bce9-b22958da49b9 · outbound

This paper cites A Priority Map for Vision-and-Language Navigation with Trajectory Plans and Feature-Location Cues.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation A Priority Map for Vision-and-Language Navigation with Trajectory Plans and Feature-Location Cues

Reference 2022

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T21:36:19.775180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.618642Z digest=sha256:69d3492e737236c5319d00aa44e942d1e48b3eb5ff86d40124aefa021a0d86cb

Observation 7e289a04-ca3e-43e1-80e6-5d0bec598cc0 · outbound

This paper cites Available: https: //api.semanticscholar.org/CorpusID: 256390509.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Available: https: //api.semanticscholar.org/CorpusID: 256390509

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.120847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T21:36:19.579993Z digest=sha256:14172197f2708efb984d86eec357da5aa28f6c197fbcf39e4138b9a50b02efb5

Pith citing papers

Observation 259dd385-779b-411d-8b2e-69fc3bc10033 · inbound

LogisticsVLN: Vision-Language Navigation For Low-Altitude Terminal Delivery Based on Agentic UAVs cites this paper.

LogisticsVLN: Vision-Language Navigation For Low-Altitude Terminal Delivery Based on Agentic UAVs NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:55:39.925157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:55:39.925157Z digest=sha256:6549bcf69ce8aad65e757e24f8671307fd2b739c66949d6b377d399fa9bd3ba5

Observation ae079907-e5e0-440d-8ecd-307e18499d80 · inbound

UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents cites this paper.

UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:11.373098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:15:11.373098Z digest=sha256:9c7da2fea4bdb14d2ba23d1354a63d4411962427cee2c35374eb46f210450a37

Observation f85f8500-c1cd-416f-ab69-baf24b23d718 · inbound

When Large Language Models Meet UAV Projects: An Empirical Study from Developers' Perspective cites this paper.

When Large Language Models Meet UAV Projects: An Empirical Study from Developers' Perspective NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T16:33:25.101602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:33:25.101602Z digest=sha256:97ec2d0a226e6b53a56a84e4849d763708e5031a5761a30e0057c4f4c50b0760

Observation ca3e084c-5b35-4cf2-ad75-2483fd4f2113 · inbound

Fly0: Persistent Metric Anchoring for Zero-Shot Aerial Vision-Language Navigation cites this paper.

Fly0: Persistent Metric Anchoring for Zero-Shot Aerial Vision-Language Navigation NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T05:36:30.808418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:36:30.808418Z digest=sha256:2d054fee3e23b217c7bdb243de1ae0c6cd84e4785948277aedff757f45f8aa3d

Observation e74eac90-090b-4076-b615-6751be7e3df7 · inbound

Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models cites this paper.

Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.037911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T18:20:12.463846Z digest=sha256:916fc309e6607ef89182bd25322fb5ea0d8eadfc7711e6485b75c6c4fc29364d

Observation 25d503f7-db0c-4a74-97b7-392f4d148dcc · inbound

How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace cites this paper.

How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:25:57.188006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T18:09:39.510770Z digest=sha256:c571498261419ee4a76df789dcf001557c5a9c5f0008efcd90aaddc8ce17651c

Observation 60634c23-66a7-4e1e-a62a-88033472dc05 · inbound

HTNav: A Hybrid Navigation Framework with Tiered Structure for Urban Aerial Vision-and-Language Navigation cites this paper.

HTNav: A Hybrid Navigation Framework with Tiered Structure for Urban Aerial Vision-and-Language Navigation NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.559653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T18:19:00.107417Z digest=sha256:928ca41cb816efd6abd70ee939e069f5de279833079d1e34e8c070b2c427a83d

Observation 281c3450-e099-4200-87e8-6ea17bff0105 · inbound

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap cites this paper.

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:29.748206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T13:48:08.135538Z digest=sha256:32ee7272b17dd27265b49b270f603b6947b55999068543af673e069c8d47e5a5

Observation 507a31ba-0e73-43e3-8858-b22ed6148c8f · inbound

FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation cites this paper.

FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:17:37.797250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T08:13:12.029402Z digest=sha256:a46bf0d383f563eb9596c383ce0fdcaad2f90e06f43c0dbd0f5dd0816c0df9b8

Observation d6299eff-d61f-4149-acc2-8a2142a33336 · inbound

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification cites this paper.

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:56:47.412900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T06:55:33.464951Z digest=sha256:d26e213ac66a09b1bd948d7e505f48c2c5304ad06bc7423798e9f13a6209d1ab

Observation 07109b6f-f89c-44c0-b6d8-90af4a1690df · inbound

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification cites this paper.

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-12T19:10:41.883296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:10:41.883296Z digest=sha256:85804bf55f4b5711af348a932e2c958a6a024dc8755da0f562f3894129651f7c

Observation 910f37db-f518-45b3-9204-338162d99e22 · inbound

PEACE: A Planner-Executor Agent with Constraint Enforcement for UAVs cites this paper.

PEACE: A Planner-Executor Agent with Constraint Enforcement for UAVs NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:53:40.407954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T16:53:17.504094Z digest=sha256:a17db9d796a861b76b889ac612d0688a55dfa622a74ed2acbfb121279680fe5d

Observation cd4be5a4-5079-4ea6-a506-b9d38a0f776d · inbound

See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View cites this paper.

See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:19:30.526950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T18:15:12.176935Z digest=sha256:bba32bb7199e34a1630dabda83a35471e93739cb7b1765f534d184241cb182d5

Observation 2671ee09-a933-488d-acac-8753d74c1d37 · inbound

AeroVerse-SatAgent: UAV-Satellite Collaborative Spatial Reasoning Inspired by the Dual Visual Pathway Theory of Cognitive Neuroscience cites this paper.

AeroVerse-SatAgent: UAV-Satellite Collaborative Spatial Reasoning Inspired by the Dual Visual Pathway Theory of Cognitive Neuroscience NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:45:39.825514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-01T06:18:35.494240Z digest=sha256:a84abd25850531b8be63d0721516b5a439866b09501d408802fb3917b3c677d4

Observation 246d65ed-a381-499f-8135-741a3e9cd573 · inbound

DynFly: Dynamic-Aware Continuous Trajectory Generation for UAV Vision-Language Navigation in Urban Environments cites this paper.

DynFly: Dynamic-Aware Continuous Trajectory Generation for UAV Vision-Language Navigation in Urban Environments NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:41.213953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-01T05:35:32.213470Z digest=sha256:64bde3502b97cd6941ea1005bc623598e1a2a407e78736cf114e621658e538d5

Observation aba18c66-6fb9-4f43-a8c9-1ec819251b22 · inbound

DynFly: Dynamic-Aware Continuous Trajectory Generation for UAV Vision-Language Navigation in Urban Environments cites this paper.

DynFly: Dynamic-Aware Continuous Trajectory Generation for UAV Vision-Language Navigation in Urban Environments NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:18:59.686028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-03T22:11:44.933463Z digest=sha256:c6096c152a1b452c3b5e143dbb855f2a8b81c7f71681aa3b35e57b533577135b

Observation 82cfd643-ee91-4fea-a522-1a1889a1a7ad · inbound

Towards Effcient Low Altitude Sensing: A Dual Heterogeneous Graph Learning Method for UAV Task Allocation cites this paper.

Towards Effcient Low Altitude Sensing: A Dual Heterogeneous Graph Learning Method for UAV Task Allocation NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-11T20:35:41.043764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:35:41.043764Z digest=sha256:8f3282fd74f6fbfa548988bfe214f839cbe26e0b29ab572d3971b41d26efbbd9

Observation 98a73178-a137-4fd4-ad4f-d5c0ced63f36 · inbound

A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation cites this paper.

A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 263

Resolution
unresolved
no resolver link, observed 2026-07-14T15:39:59.878169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:39:59.878169Z digest=sha256:a3df31cbf94f9450830b5b9ebfe8c8fa90b18c74d9413010b4768e82793ff2ad

Observation c73e5cef-efe4-4e20-998a-fa9370c7a9d2 · inbound

Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory cites this paper.

Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T04:58:23.085764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:58:23.085764Z digest=sha256:1906487d792a793c6e54fdaee00aa07f8a128bd6da0e911f58c0bedd4581b87f

Observation 8694a104-4ed1-4bc4-994e-e90d4ff8b656 · inbound

Deliberate Before You Fly: Vision-Guided Spatial Deliberation for UAV See-and-Reach Navigation cites this paper.

Deliberate Before You Fly: Vision-Guided Spatial Deliberation for UAV See-and-Reach Navigation NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:59.145803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:42:59.145803Z digest=sha256:5e828de7ead7d976456aaf22cb06cbbea5ccb7a31ec858849b171b643873f37d