Pith. sign in

REVIEW 5 cited by

Aerial Vision-and-Dialog Navigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.12219 v3 pith:TJGPDLIV submitted 2022-05-24 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords navigationdroneaerialattentionavdndatasetfollowershuman
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The ability to converse with humans and follow natural language commands is crucial for intelligent unmanned aerial vehicles (a.k.a. drones). It can relieve people's burden of holding a controller all the time, allow multitasking, and make drone control more accessible for people with disabilities or with their hands occupied. To this end, we introduce Aerial Vision-and-Dialog Navigation (AVDN), to navigate a drone via natural language conversation. We build a drone simulator with a continuous photorealistic environment and collect a new AVDN dataset of over 3k recorded navigation trajectories with asynchronous human-human dialogs between commanders and followers. The commander provides initial navigation instruction and further guidance by request, while the follower navigates the drone in the simulator and asks questions when needed. During data collection, followers' attention on the drone's visual observation is also recorded. Based on the AVDN dataset, we study the tasks of aerial navigation from (full) dialog history and propose an effective Human Attention Aided Transformer model (HAA-Transformer), which learns to predict both navigation waypoints and human attention.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs

    cs.CV 2026-07 conditional novelty 6.0 of 10

    UAV-DualCog evaluates multimodal LLMs on self-state and environment-state reasoning in aerial images and videos, and shows current models are unreliable at spatial and temporal grounding.

  2. Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent

    cs.RO 2025-06 conditional novelty 5.0 of 10

    An open-source ROS2/PX4 framework using locally hosted LLMs and VLMs enables natural language drone commands, with the best simulated mission success rate at 40%.

  3. Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A dual-branch text-imagination system, with one LLM branch for state estimation and one for candidate-direction description, improves R2R navigation success over prior LLM-based VLN methods.

  4. UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation

    cs.RO 2025-01 conditional novelty 5.0 of 10

    UAV-VLA generates drone flight plans from natural language using satellite imagery, GPT, and Molmo, and introduces a 30-image benchmark, but its evaluation against a single human operator is weak.

  5. No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation

    cs.CV 2026-07 conditional novelty 4.0 of 10

    Test-time scaling—parallel candidate generation, iterative self-correction, and multi-criteria selection—improves a frozen UAV navigation VLM's success rate by about 2 percentage points on the TravelUAV benchmark.

Pith tools