Pith. sign in

Paper Citation Record · LEDGER

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

As of 20 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 20 inbound Pith citation observations for arXiv:2411.08579.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.08579 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:36:19.687871Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:55:39.925157Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:19:30.523454Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact1
  • verified fuzzy43
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 04fe7347-d2af-4c67-8b3f-6876f353ca2e · outbound

This paper cites Aerialvln: Vision-and-language navigation for uavs,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Aerialvln: Vision-and-language navigation for uavs,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.288573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.524424Z digest=sha256:5e508aa1a97afb6f796f714de3dc596594515eda9f29c38ed003432ac0cb2e7c

Observation d02febe6-863d-4560-b23a-d4476a0d7b9c · outbound

This paper cites Vision- and-language navigation: Interpreting visually-grounded navigation in- structions in real environments,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Vision- and-language navigation: Interpreting visually-grounded navigation in- structions in real environments,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.280133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.527708Z digest=sha256:1ec4a4b0a2f8609549be5669ecc7cddd908fcaa1ee9a0989867da47af54141ca

Observation 50c515e4-f099-49ca-88c6-f8938fb987ec · outbound

This paper cites Reverie: Remote embodied visual referring expression in real indoor environ- ments,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Reverie: Remote embodied visual referring expression in real indoor environ- ments,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.271809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.530423Z digest=sha256:7b3887315f4000fb50541f56871866966f31bb1545599647b21d58696406a018

Observation c028d0b6-99e3-438b-9115-b29d8be19baf · outbound

This paper cites Stay on the path: Instruction fidelity in vision-and-language navigation,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Stay on the path: Instruction fidelity in vision-and-language navigation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.263047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.533221Z digest=sha256:f6207240bdc83c1388fd4a2992341398898d6ef2251fd20bf666461ec2232ab9

Observation 8167c85e-b03d-4658-8d39-7bb3a5ee43ef · outbound

This paper cites Room-across- room: Multilingual vision-and-language navigation with dense spa- tiotemporal grounding,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Room-across- room: Multilingual vision-and-language navigation with dense spa- tiotemporal grounding,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.253757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.536142Z digest=sha256:47d9c0ff3a6d697bfac29b3685f149094578866af18b3d58f6f41502249a494a

Observation 481fcf3f-a758-4e06-946d-c1609a40be6d · outbound

This paper cites Beyond the nav-graph: Vision-and-language navigation in continuous environments,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Beyond the nav-graph: Vision-and-language navigation in continuous environments,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.246465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.538936Z digest=sha256:83af2e69b0ef5cca08c839383f80082cb3c54fa91fdcfc66780f45e67cc3e1f1

Observation 5ff2afa4-1dda-4515-b01d-6d3af43202b5 · outbound

This paper cites Touchdown: Natural language navigation and spatial reasoning in visual street environments,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Touchdown: Natural language navigation and spatial reasoning in visual street environments,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.239046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.541736Z digest=sha256:0343fe9683394fcd175e6beef7f7bbe2497932f91e993f7507bd8c3839ecaca4

Observation 6b3e3a5c-0025-4063-8dcb-9ee6423bebfa · outbound

This paper cites Dji drone solutions for inspection and infrastructure construction in the oil and gas industry,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Dji drone solutions for inspection and infrastructure construction in the oil and gas industry,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.229332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.544208Z digest=sha256:ebb1d66ea5044b6b15e1345e82f21d8d69595d4069f3c2fca42f72acfd183cbb

Observation 8558fc49-85bb-474c-b317-86724bd9dc96 · outbound

This paper cites Dji drone solutions for optimizing operations in the public safety industry,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Dji drone solutions for optimizing operations in the public safety industry,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.220369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.546761Z digest=sha256:ddfced4032513187a4c846521613e20dad294c6d16bbc855799087656228ab54

Observation 55091f4a-2bd7-44c1-964f-5fe7f245fcef · outbound

This paper cites Dji drone solutions for surveying, urban planning, aec, and natural resource management,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Dji drone solutions for surveying, urban planning, aec, and natural resource management,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.211167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.549505Z digest=sha256:3cbcd121b9cf2f318a758e568f1ac16fff0467314db6f5ac549b9b630506e884

Observation 9bc98473-640a-4dbb-99d8-4343722935b2 · outbound

This paper cites Hop+: History- enhanced and order-aware pre-training for vision-and-language naviga- tion,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Hop+: History- enhanced and order-aware pre-training for vision-and-language naviga- tion,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.202399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.551791Z digest=sha256:6b96a69646ab681fdaeac4732f7ed7263c67c48c2b46131a29e7f0585fcba966

Observation 14b69d68-b0b4-43a7-b915-71a6246b0d82 · outbound

This paper cites Correctable landmark discovery via large models for vision-language navigation,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Correctable landmark discovery via large models for vision-language navigation,

Reference 12

Resolution
metadata mismatch
raw_fallback, observed 2026-08-12T21:36:19.880078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.554070Z digest=sha256:83ee50f4dcd29ca1ef92cd6d3734ca5750bd857f0df965ac4c9ed0733475e648

Observation 05b6b6d9-d36c-44ca-bdc5-ad72ac5e78dc · outbound

This paper cites Learning to follow and generate instructions for language-capable navigation,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Learning to follow and generate instructions for language-capable navigation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.193474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.556607Z digest=sha256:673d732743b3e73d858d2b9b3ed24f113bf791068c2505a092b8791b57f92e73

Observation ac166ce2-4757-4d7c-b6e3-40c2a72d9b54 · outbound

This paper cites ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.559058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.559058Z digest=sha256:8bc8cb2c6d7b17997eb914a9e70ad46b1210d6a2b86e4a6e9e03c84f9895bc6b

Observation be234d5e-41ea-4056-acee-55b1c92de0d8 · outbound

This paper cites Towards deviation-robust agent navigation via perturbation-aware con- trastive learning,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Towards deviation-robust agent navigation via perturbation-aware con- trastive learning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.183382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.562158Z digest=sha256:ccf4c8086ef60c41edae63ff7602943322c8bcf9cf6c200445913f7b5c3b6182

Observation ed55ea5f-55fd-443b-ad70-5c8664cf7d43 · outbound

This paper cites A survey on vision-based uav navigation,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation A survey on vision-based uav navigation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.173548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.564603Z digest=sha256:eec6a12e6ad64ed46f56ecbefa207c0b31504a79f1cb0d0c2dc956485bae3cfb

Observation 8b95d68b-795b-46f2-977f-e43f83ed39ee · outbound

This paper cites Vision-based navigation of unmanned aerial vehicles,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Vision-based navigation of unmanned aerial vehicles,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.163859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.567166Z digest=sha256:a451ed030360170f6def62176caea79c82b4759668b181a25ddc206ad0132f41

Observation 87158f5f-f303-41c3-9aae-244024c4b5cc · outbound

This paper cites Mapping instructions to actions in 3d environments with visual goal prediction,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Mapping instructions to actions in 3d environments with visual goal prediction,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.154605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.569665Z digest=sha256:5f2ec1173226518672b5438e77abcfeebc4e71fd91c2cd69686551af69d35b12

Observation 3dfead87-f1b0-423d-b421-8bc819ef90a1 · outbound

This paper cites Following high-level navigation instructions on a simulated quadcopter with imitation learning,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Following high-level navigation instructions on a simulated quadcopter with imitation learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.145386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.572086Z digest=sha256:8e7c822928125dfc88d441f4560287444a888593e1598d958e05a7d1087d7656

Observation 740247c1-d6d0-4488-acc2-50780c0f342b · outbound

This paper cites Mapping naviga- tion instructions to continuous control actions with position-visitation prediction,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Mapping naviga- tion instructions to continuous control actions with position-visitation prediction,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.135601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.574502Z digest=sha256:ccb77da3d471658c5bff9fd85da4a29200155728c8f8cf2d717bb81a13ed3220

Observation 5c176195-5011-4465-afe1-9ca700386462 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.577236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.577236Z digest=sha256:438db5a342bfd6062ed93f473154faadef85b67cdb5f4a55724f2d422b472119

Observation 08c75f42-d292-497b-aa46-88e384b5241a · outbound

This paper cites Visual Instruction Tuning.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Visual Instruction Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.582527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.582527Z digest=sha256:61ddf193e465965bd3b4f0307f52c1e679f7b711465ea091545f0792b8e370ca

Observation 8081e15b-0bed-4e6a-8e0f-a7ef0b1f98f4 · outbound

This paper cites Grounded language-image pre- training,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Grounded language-image pre- training,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.111581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.585270Z digest=sha256:e73c70b6eb2b7c711a3a93b8a20ebb178f162e11e0d4f0e9254549c317a72764

Observation d5640189-f98a-43fb-ad4d-01639967cd94 · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Instructblip: Towards general-purpose vision- language models with instruction tuning,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.587760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.587760Z digest=sha256:eb4084747f7a1684077724afecabcfacbd303dfcd470875b47d0cbd3f07159b3

Observation b3ff0a49-d507-4da7-bbe7-e2ccbd7c251b · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.590227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.590227Z digest=sha256:7aa5815e9dab17b89fd3170fd554e308dbd6841461fed6cba50ac82d8a2147c4

Observation cd3e0007-e2bd-4113-9be9-ea1546bf5de0 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.592961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.592961Z digest=sha256:e7c5afd9182e8194f892e219bd9bc531fcaa66929c1f24d9497c4e42059a3471

Observation e42cd9e9-2987-4fd5-b6fe-175616d3c02d · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.595607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.595607Z digest=sha256:1dd90148b5c79a017d6a0f302e15835ebc54f7c7024fedae92d6fce59f70ce23

Observation e700f938-6099-49c1-9cc8-9ad87223acd7 · outbound

This paper cites Qwen Technical Report.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Qwen Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.598472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.598472Z digest=sha256:64921dc26f603c1e423c326858bba736a1db10f4322c9435ae300e9d68072e3b

Observation 361d3cf7-3a01-4b60-aaca-0466c615004e · outbound

This paper cites Gsv-cities: Toward appropriate supervised visual place recognition,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Gsv-cities: Toward appropriate supervised visual place recognition,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.097195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.601061Z digest=sha256:9afa7e1009cc8c47b7d58382831468dd6b5f2dbed55608c689eaaba5e3f94fa0

Observation 1faea57a-3bfb-4221-9c99-11a2f445dcb8 · outbound

This paper cites The streetlearn environment and dataset,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation The streetlearn environment and dataset,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.088346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.603333Z digest=sha256:c12aef37b5b07e671aedb460ca2081ccf44363c612a0678959c2f646ec5500af

Observation 5c2dcecb-229a-40f0-802b-a76f9f343137 · outbound

This paper cites Learning to follow directions in street view,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Learning to follow directions in street view,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.079866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.605868Z digest=sha256:e12e93ffbeb77692ade8cd9c03e4b830f62ff6fe6ca105fc2b074f66ad8b556e

Observation a88d4c06-7bf1-4479-ba47-f36f1c8985cd · outbound

This paper cites Retouchdown: Adding touchdown to streetlearn as a shareable resource for language grounding tasks in street view,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Retouchdown: Adding touchdown to streetlearn as a shareable resource for language grounding tasks in street view,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.070836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.608453Z digest=sha256:f3256cae80aab1de17432fa306761920f1ae3b2c9377d07454fb51a2d868d66d

Observation 27ec9b46-ea08-4b44-912e-985f984d3931 · outbound

This paper cites Silg: The multi-environment symbolic interactive language grounding benchmark,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Silg: The multi-environment symbolic interactive language grounding benchmark,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.061931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.610952Z digest=sha256:a7839a7cf49c67cf54535901698b5234cdd0c16c059ad3a20604fccd24e4326c

Observation cb78d08c-45a3-49cd-b72d-c196970336c2 · outbound

This paper cites Outdoor vision-and-language navigation needs object-level alignment,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Outdoor vision-and-language navigation needs object-level alignment,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.053117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.613415Z digest=sha256:1e4a1684eeb9417100cb73aed5b1e0d366b99b8a0dd35505b35befcf1e40d9fe

Observation ebeb6a68-5332-4048-95e9-2b7fd5c47faa · outbound

This paper cites A priority map for vision-and- language navigation with trajectory plans and feature-location cues,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation A priority map for vision-and- language navigation with trajectory plans and feature-location cues,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.044234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.615811Z digest=sha256:95c1b9aa1dfb3aabc8f47d997f4589b7a0b9547c6075f3c4fab6f39a57dfccc2

Observation c4797cf4-cd03-4f72-8c2d-85c8e909d761 · outbound

This paper cites Multimodal text style transfer for outdoor vision- and-language navigation,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Multimodal text style transfer for outdoor vision- and-language navigation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.035193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.621791Z digest=sha256:0c22ada9fe76f7cf2f61d4bf17a668943201056df012cfea3424212c5e503ac0

Observation 7a35650c-9e21-4afa-884b-6f1c7c6c42bf · outbound

This paper cites LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.624438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.624438Z digest=sha256:6655c74a265e7242538ef16cccf274b772f3b05bf480f3b4120b368c9ab610f0

Observation 9a42f1b5-c8e9-4924-98db-5bf2536bd6b2 · outbound

This paper cites VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.627503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.627503Z digest=sha256:ed9dda1348dafa568b0a7e30787cc3bf112a76068ad656941848a820969b29df

Observation 3f7c5d22-0c64-4ab6-9eea-8ee5d29860b3 · outbound

This paper cites Opt: Open pre-trained transformer language models,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Opt: Open pre-trained transformer language models,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.025980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.630225Z digest=sha256:3e7f0e40c3aefc7593e1b07b2412a348732a054427f92d4fef0f167cd184888b

Observation b98990a4-4241-4958-a765-dd272f545f0e · outbound

This paper cites Palm-e: An embodied multimodal language model,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Palm-e: An embodied multimodal language model,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.017244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.632744Z digest=sha256:ffa803b7df727d983c1a061fe612cd060e284e7447fb2861a0f4905849263212

Observation 671d8ebc-8939-401d-ae97-9604757d5c70 · outbound

This paper cites Training language models to follow instructions with human feedback,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Training language models to follow instructions with human feedback,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.008623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.635530Z digest=sha256:ae85bee73e028b93e13d6e750bb2c73f5868427f069ddcda206d9850d60770f6

Observation 68054db1-e704-4b2a-902d-dd2bbbaa29a2 · outbound

This paper cites GPT-4 Technical Report.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation GPT-4 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.638214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.638214Z digest=sha256:bbee7accac2d5a06e57f22a9bf06b0e8c5a12b8d12ad760b708a4db8580bdb92

Observation e30296cb-10a5-4f4f-9065-a72edb665336 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation LLaMA: Open and Efficient Foundation Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.641203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.641203Z digest=sha256:23bdc39d04479726b4ac03d71c61d9568e0b307a02d78197c3963f0b39f22e66

Observation e59f452c-8b74-461c-aecf-fbac03409379 · outbound

This paper cites ESC: Exploration with Soft Commonsense Constraints for Zero-shot Object Navigation.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation ESC: Exploration with Soft Commonsense Constraints for Zero-shot Object Navigation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.643739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.643739Z digest=sha256:f87c9ba9fc18a33e7daea6010dfe5af572c9e0922381409e18e4cdd4c6e31977

Observation 018b6d5a-777a-4d61-b9be-40de66b54966 · outbound

This paper cites NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.646539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.646539Z digest=sha256:1c59d17efa76f5342e9e0d506fc0a6f56d60d78ea31ccefc449ad35cd2c4e064

Observation 981a5aad-1fb3-4393-b193-a96efc1b2a7b · outbound

This paper cites Neural slam: Learning to explore with external memory,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Neural slam: Learning to explore with external memory,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.999617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.649143Z digest=sha256:e3c4338f0b4d0d0696d9d9faab9366571a8f0168e616fea439daa5ae8fadb434

Observation 97d9714a-e0d8-40ed-a4a6-f453f5a2b1c9 · outbound

This paper cites Egomap: Pro- jective mapping and structured egocentric memory for deep rl,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Egomap: Pro- jective mapping and structured egocentric memory for deep rl,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.990743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.651661Z digest=sha256:5732f04a1720af25c558d961a4a090f2d9eddb5f81b3c8ebaf560c80358ba258

Observation f0dd38f3-aeee-4319-b76c-8bc758bef865 · outbound

This paper cites Semantic mapnet: Building allocentric semanticmaps and representations from egocentric views,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Semantic mapnet: Building allocentric semanticmaps and representations from egocentric views,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.981765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.654249Z digest=sha256:0e6dc7170ab569694e022544aa2729d567b25a469661a1626c10ee8e300653f3

Observation 5752096d-59e7-40e4-a8be-53be577cebd1 · outbound

This paper cites Audio visual language maps forrobot navigation,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Audio visual language maps forrobot navigation,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.972719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.656831Z digest=sha256:9e8035c3f9df5ec3dc816cef18a2787b5bf88dbd08eae770c8d604881a18c9be

Observation d0e7dad1-cedc-4b88-8ddc-f46816f24fc3 · outbound

This paper cites Mapnet: An allocentric spatial memory for mapping environments,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Mapnet: An allocentric spatial memory for mapping environments,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.963827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.659298Z digest=sha256:8b2b7bfdad3730f8589fe3f723eac4a72135503cdbc9fe91829a3995ebb26258

Observation 02cfc938-f6e9-4897-98b1-d6ee4feea425 · outbound

This paper cites BEVBert: Multimodal Map Pre-training for Language-guided Navigation.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation BEVBert: Multimodal Map Pre-training for Language-guided Navigation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.661906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.661906Z digest=sha256:b3338969fe5bc47ed7f7eaf289116ebbfaab91f4db467e7e3c7454da15872905

Observation 9b204acb-cc5b-434c-b783-d05be3d74076 · outbound

This paper cites Cognitive mapping and planning for visual navigation,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Cognitive mapping and planning for visual navigation,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.954209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.664626Z digest=sha256:286fa7d9b1e4f56cd634e362d4a2bdc7e668e10d523a07222a22a00f86423398

Observation 031e2c47-b4a9-4626-9b0f-3ffd4591cdbc · outbound

This paper cites Semantic mapnet: Building allocentric semantic maps and representations from egocentric views,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Semantic mapnet: Building allocentric semantic maps and representations from egocentric views,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.943942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.667053Z digest=sha256:788feccb592fbb511351be5487c768bcb39c8cbc205d87125a1072cb782c6074

Observation 041a2fff-21b9-470c-b623-41433d2fff10 · outbound

This paper cites Cross-modal map learning for vision and language navigation,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Cross-modal map learning for vision and language navigation,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.934773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.669525Z digest=sha256:87fe8e738dc7ab4344ab80fa5417cea6b23c63367606cd52dacb9a32971537bf

Observation 569d789a-e2f9-4554-bfe0-ed6486ca583b · outbound

This paper cites Topo- logical planning with transformers for vision-and-language navigation,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Topo- logical planning with transformers for vision-and-language navigation,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.925391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.671919Z digest=sha256:f8d4fe334bea410e9489bfb44ead5fb4da3a63cb9d326323a91eafbe0a1e2ed5

Observation c54bd769-1044-4b50-b5d9-8f4b8bfd62f0 · outbound

This paper cites Generating landmark navigation instruc- tions from maps as a graph-to-text problem,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Generating landmark navigation instruc- tions from maps as a graph-to-text problem,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.915520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.674472Z digest=sha256:5a8622a9fbaedcbb4fcf66eb5521b7591995335e0fe0fe5e6a277819ff0c6ee5

Observation 81e05be7-04b8-4461-8a59-f34ae2111060 · outbound

This paper cites Microsoft COCO: Common Objects in Context.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Microsoft COCO: Common Objects in Context

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.676906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.676906Z digest=sha256:6bcf9a44ffa7b7757be2353e7973d7eafaef58c7d8b310faf504d6ad1f2003cc

Observation 2e33fcc0-dadf-4a77-ab13-678911f39f40 · outbound

This paper cites Dynamic head: Unifying object detection heads with attentions,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Dynamic head: Unifying object detection heads with attentions,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.906035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.679766Z digest=sha256:412380e37e2e002498cba2b97e9c9ea18a0eedbb306c68f36010485aedccc58e

Observation 0b591f49-95c5-4320-b752-ced7d0b05aa3 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T21:36:19.682360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:36:19.682360Z digest=sha256:481bcb54d359a3b39ff22ce0476f70a16dd5b70686f9f497ac102d32a1312753

Observation 97605107-7c0b-4f76-95fd-549fe7bb7d30 · outbound

This paper cites Analyzing Generalization of Vision and Language Navigation to Unseen Outdoor Areas.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Analyzing Generalization of Vision and Language Navigation to Unseen Outdoor Areas

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-12T21:36:19.715144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.687871Z digest=sha256:16113ff081307c44010eb713bda9e1f63a9e61359bacb0c1a2226f01e2b1e498

Observation e20f1ef2-14e0-4d8a-8444-24590a1c84f5 · outbound

This paper cites Available: https: //api.semanticscholar.org/CorpusID: 52967399.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Available: https: //api.semanticscholar.org/CorpusID: 52967399

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:19.890398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.685176Z digest=sha256:a4897ec497f82ef2b4b4fe44515ecb6aa1ebcb4ddabaea77a8870d47965ab498

Observation bcce2f57-31aa-4f75-bce9-b22958da49b9 · outbound

This paper cites A Priority Map for Vision-and-Language Navigation with Trajectory Plans and Feature-Location Cues.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation A Priority Map for Vision-and-Language Navigation with Trajectory Plans and Feature-Location Cues

Reference 2022

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T21:36:19.775180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.618642Z digest=sha256:69ccc85fd58adac6184d2b40b00f72c0bebc5c13a2a3630e9cc777ede0433855

Observation 7e289a04-ca3e-43e1-80e6-5d0bec598cc0 · outbound

This paper cites Available: https: //api.semanticscholar.org/CorpusID: 256390509.

NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation Available: https: //api.semanticscholar.org/CorpusID: 256390509

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:36:20.120847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T21:36:19.579993Z digest=sha256:3a2426a1714b548baad499b7f6c60994f864b436aed301fc7dfa0dbddb67c07d

Pith citing papers

Observation 259dd385-779b-411d-8b2e-69fc3bc10033 · inbound

LogisticsVLN: Vision-Language Navigation For Low-Altitude Terminal Delivery Based on Agentic UAVs cites this paper.

LogisticsVLN: Vision-Language Navigation For Low-Altitude Terminal Delivery Based on Agentic UAVs NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:55:39.925157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:55:39.925157Z digest=sha256:6549bcf69ce8aad65e757e24f8671307fd2b739c66949d6b377d399fa9bd3ba5

Observation ae079907-e5e0-440d-8ecd-307e18499d80 · inbound

UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents cites this paper.

UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:11.373098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:15:11.373098Z digest=sha256:9c7da2fea4bdb14d2ba23d1354a63d4411962427cee2c35374eb46f210450a37

Observation f85f8500-c1cd-416f-ab69-baf24b23d718 · inbound

When Large Language Models Meet UAV Projects: An Empirical Study from Developers' Perspective cites this paper.

When Large Language Models Meet UAV Projects: An Empirical Study from Developers' Perspective NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T16:33:25.101602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:33:25.101602Z digest=sha256:97ec2d0a226e6b53a56a84e4849d763708e5031a5761a30e0057c4f4c50b0760

Observation ca3e084c-5b35-4cf2-ad75-2483fd4f2113 · inbound

Fly0: Persistent Metric Anchoring for Zero-Shot Aerial Vision-Language Navigation cites this paper.

Fly0: Persistent Metric Anchoring for Zero-Shot Aerial Vision-Language Navigation NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T05:36:30.808418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:36:30.808418Z digest=sha256:2d054fee3e23b217c7bdb243de1ae0c6cd84e4785948277aedff757f45f8aa3d

Observation e74eac90-090b-4076-b615-6751be7e3df7 · inbound

Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models cites this paper.

Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.037911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T18:20:12.463846Z digest=sha256:ae417cfbe2cf24207ae3748721b8de4d93cbe25a1e9656671838dfd8015a1d2d

Observation 25d503f7-db0c-4a74-97b7-392f4d148dcc · inbound

How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace cites this paper.

How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:25:57.188006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T18:09:39.510770Z digest=sha256:dc63b070be3fcaf6478fdd403fb0d484994fac8e3bfeb7ebb8283b6ade7f471c

Observation 60634c23-66a7-4e1e-a62a-88033472dc05 · inbound

HTNav: A Hybrid Navigation Framework with Tiered Structure for Urban Aerial Vision-and-Language Navigation cites this paper.

HTNav: A Hybrid Navigation Framework with Tiered Structure for Urban Aerial Vision-and-Language Navigation NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.559653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T18:19:00.107417Z digest=sha256:f5532a737af926a767e50c70c53078421416b200ef41a0fab935f165e21ac051

Observation 281c3450-e099-4200-87e8-6ea17bff0105 · inbound

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap cites this paper.

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:29.748206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T13:48:08.135538Z digest=sha256:721546cb4e6a38b0e5334388d0d8fc49a5f85477308056f0b352e56a73378362

Observation 507a31ba-0e73-43e3-8858-b22ed6148c8f · inbound

FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation cites this paper.

FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:17:37.797250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T08:13:12.029402Z digest=sha256:561043237bb8a8a5a1343e0cb0485f24b0d4bc672c1e4003b04fe3e4b955a5c6

Observation d6299eff-d61f-4149-acc2-8a2142a33336 · inbound

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification cites this paper.

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:56:47.412900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T06:55:33.464951Z digest=sha256:0d315542e4d8e436aab1634a30656ae8c74d264a6418ca46ccaea3830de9ff85

Observation 07109b6f-f89c-44c0-b6d8-90af4a1690df · inbound

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification cites this paper.

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-12T19:10:41.883296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:10:41.883296Z digest=sha256:85804bf55f4b5711af348a932e2c958a6a024dc8755da0f562f3894129651f7c

Observation 910f37db-f518-45b3-9204-338162d99e22 · inbound

PEACE: A Planner-Executor Agent with Constraint Enforcement for UAVs cites this paper.

PEACE: A Planner-Executor Agent with Constraint Enforcement for UAVs NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:53:40.407954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T16:53:17.504094Z digest=sha256:167f12e6704be8cef9e2f5ff04cdbc6067cbdfbc26f3c7872f7b0309b7a2e147

Observation cd4be5a4-5079-4ea6-a506-b9d38a0f776d · inbound

See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View cites this paper.

See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:19:30.526950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T18:15:12.176935Z digest=sha256:b0ce0a178e3405b2941ee436604b9b60252e2908ad61a0636c2d728dcbcf3826

Observation 2671ee09-a933-488d-acac-8753d74c1d37 · inbound

AeroVerse-SatAgent: UAV-Satellite Collaborative Spatial Reasoning Inspired by the Dual Visual Pathway Theory of Cognitive Neuroscience cites this paper.

AeroVerse-SatAgent: UAV-Satellite Collaborative Spatial Reasoning Inspired by the Dual Visual Pathway Theory of Cognitive Neuroscience NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:45:39.825514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T06:18:35.494240Z digest=sha256:623debbbf7f5175652ad72dc3dfb888ec00af14bb76796e01b418e403d28d4e4

Observation 246d65ed-a381-499f-8135-741a3e9cd573 · inbound

DynFly: Dynamic-Aware Continuous Trajectory Generation for UAV Vision-Language Navigation in Urban Environments cites this paper.

DynFly: Dynamic-Aware Continuous Trajectory Generation for UAV Vision-Language Navigation in Urban Environments NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:41.213953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T05:35:32.213470Z digest=sha256:b3efa4140e3fb1f3fbef39678cb09fabe29a01e3b9a269effc0982274ac04965

Observation aba18c66-6fb9-4f43-a8c9-1ec819251b22 · inbound

DynFly: Dynamic-Aware Continuous Trajectory Generation for UAV Vision-Language Navigation in Urban Environments cites this paper.

DynFly: Dynamic-Aware Continuous Trajectory Generation for UAV Vision-Language Navigation in Urban Environments NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:18:59.686028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T22:11:44.933463Z digest=sha256:b37084300dd7179383ed04f1d5467ecd19f260077d3098fce36b3f0d41d22423

Observation 82cfd643-ee91-4fea-a522-1a1889a1a7ad · inbound

Towards Effcient Low Altitude Sensing: A Dual Heterogeneous Graph Learning Method for UAV Task Allocation cites this paper.

Towards Effcient Low Altitude Sensing: A Dual Heterogeneous Graph Learning Method for UAV Task Allocation NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-11T20:35:41.043764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:35:41.043764Z digest=sha256:8f3282fd74f6fbfa548988bfe214f839cbe26e0b29ab572d3971b41d26efbbd9

Observation 98a73178-a137-4fd4-ad4f-d5c0ced63f36 · inbound

A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation cites this paper.

A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 263

Resolution
unresolved
no resolver link, observed 2026-07-14T15:39:59.878169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:39:59.878169Z digest=sha256:a3df31cbf94f9450830b5b9ebfe8c8fa90b18c74d9413010b4768e82793ff2ad

Observation c73e5cef-efe4-4e20-998a-fa9370c7a9d2 · inbound

Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory cites this paper.

Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T04:58:23.085764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:58:23.085764Z digest=sha256:1906487d792a793c6e54fdaee00aa07f8a128bd6da0e911f58c0bedd4581b87f

Observation 8694a104-4ed1-4bc4-994e-e90d4ff8b656 · inbound

Deliberate Before You Fly: Vision-Guided Spatial Deliberation for UAV See-and-Reach Navigation cites this paper.

Deliberate Before You Fly: Vision-Guided Spatial Deliberation for UAV See-and-Reach Navigation NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:59.145803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:42:59.145803Z digest=sha256:5e828de7ead7d976456aaf22cb06cbbea5ccb7a31ec858849b171b643873f37d