Pith. sign in

REVIEW 4 major objections 5 minor 2 cited by

doScenes: An Autonomous Driving Dataset with Natural Language Instruction for Human Interaction and Vision-Language Navigation

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper claims that doScenes is the first public real-world dataset to pair autonomous-driving sensor clips with natural-language driving instructions and referentiality tags, linking imperative language to observed vehicle motion.

desk verdict Useful nuScenes instruction layer, but the 'first real-world' claim ignores Talk2Car and the annotation stats are unverified. read the letter →

arxiv 2412.05893 v1 pith:GMXI7UM2 submitted 2024-12-08 cs.CV cs.AI

classification cs.CVcs.AI
keywords autonomousdrivingnaturallanguageinstructionvision-languagenavigationdatasetannotationreferentialitymotionplanningnuSceneshuman-vehicleinteraction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces doScenes, a dataset built by retroactively annotating 1,000 real-world driving clips from nuScenes with short, passenger-style natural-language instructions. Each clip receives the instruction that would have caused the observed motion, plus a tag indicating whether the instruction refers to static objects, dynamic objects, both, or neither. The authors argue this is the first public real-world dataset to connect imperative language directly to vehicle motion, in contrast to simulated instruction datasets and real-world datasets that only caption risks or scenes. If the annotations are reliable, doScenes would let vision-language-action models learn instruction-conditioned motion planning from real sensor data rather than from fixed command sets.

What carries the argument

The central mechanism is the retroactive taxi-test annotation procedure: for each nuScenes clip, five independent annotators imagine riding as a passenger and write the short instruction that would initiate the observed motion, optionally with multiple phrasings per scene. The referentiality tags (none, static, dynamic, both) operationalize whether the instruction requires continued observation of scene objects, which distinguishes closed-loop evaluation from open-loop prediction. This procedure turns an existing large multimodal driving corpus into instruction-conditioned training data without instrumented in-vehicle recording.

What would settle it

Collect real in-car passenger instructions synchronized with ego trajectories and compare their content and referentiality distribution to the taxi-test annotations; systematic differences would undermine the proxy. A simpler check is measuring inter-annotator agreement on the taxi test for a sample of scenes, since near-chance agreement would show the causal instruction is not reliably recoverable from playback.

Watch

Extended reading notes

Core claim

doScenes augments all 1,000 nuScenes scenes with retroactive driving-instruction annotations and referentiality tags. Annotators watching sensor playback apply the taxi test, asking what instruction a passenger would give to trigger the observed maneuver, and mark whether the instruction is non-referential, static-referential, dynamic-referential, or both. The dataset statistics report 535 non-referential, 214 static-referential, 159 dynamic-referential, and 93 both-referential annotations. The paper claims this makes doScenes the first public real-world dataset to provide driving instructions and referentiality information as the natural-language annotation, creating a link between imperative language and motion for autonomous vehicles, and that it supports nuanced, flexible responses beyond the predefined action sets used in simulated instruction datasets.

Load-bearing premise

The dataset's value rests on the assumption that instructions written after the fact by annotators watching sensor playback faithfully capture the instruction a real passenger would have given to produce the observed driving.

Editorial extensions

If this is right

  • Vision-language-action models trained on doScenes could respond to open-vocabulary natural-language commands rather than a fixed set of predefined directives.
  • The static versus dynamic referentiality tags allow training and evaluation on subsets where models must either ground instructions in map-visible static objects or track moving objects that require closed-loop observation.
  • The annotations support the reverse task of generating a natural-language description of a trajectory, contributing to interpretable and interactive autonomous driving.
  • Multi-stage motion planning can be studied, since some 12-second clips contain maneuvers longer than a single instruction, and accurate response may appear only in the first portion of a path.
  • Models using only rasterized maps may handle non-referential instructions, but referential instructions will require LiDAR or front-view images, giving a natural test bed for sensor-modality choices.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the taxi-test proxy holds, doScenes could become a standard benchmark for instruction-conditioned trajectory prediction, letting researchers measure how much language grounding improves forecasting beyond scene-only models.
  • A testable extension is to compare model performance on static-referential versus dynamic-referential instructions, predicting that dynamic ones are harder because the referred object's future state is what drives the motion plan.
  • The multiple annotations per scene capture paraphrased instructions with similar intent, which could be used to evaluate how robust language-grounding models are to varied phrasings of the same directive.
  • If retroactive annotation proves reliable, it offers a low-cost route to convert other existing driving corpora into instruction datasets, accelerating data collection without instrumented vehicles.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces doScenes, a dataset built by retroactively annotating all 1,000 nuScenes 12-second driving scenes with natural language instructions and referentiality tags (non-referential, static-referential, dynamic-referential, and both). Five annotators used a 'taxi test' heuristic to infer what passenger instruction, if any, would have caused the observed ego motion. The paper reports aggregate statistics on referentiality, discusses the retroactive-annotation limitation, and proposes applications in vision-language navigation and vision-language-action models. The dataset is made public through a GitHub link.

Significance. If the annotations are reliable, doScenes could provide a useful resource for linking imperative language to real-world driving motion, complementing simulation-based instruction datasets and enabling research on referential instructions that require closed-loop observation. The explicit static/dynamic referentiality tags are a useful and relatively novel annotation axis. However, the paper does not yet provide the evidence needed to assess annotation reliability, and its central novelty claim is overstated by the omission of Talk2Car, a public real-world dataset that already supplies natural-language driving commands, referred-object annotations, and ego trajectories on nuScenes. The retroactive-annotation approach is honestly acknowledged, but its validity for establishing an instruction-to-motion link is left unvalidated.

major comments (4)
  1. [II-C] The claim that doScenes is 'the first public real-world dataset to provide driving instructions and referentiality information as the natural language annotation' is not defensible, because Talk2Car (Deruyttere et al., 2019) already provides natural-language driving commands, referred-object annotations, and ego trajectories on nuScenes data. The absence of any citation or discussion of Talk2Car in Section II is a substantive omission that directly affects the paper's central contribution claim.
  2. [III] The annotation statistics in Table II sum to 1,001 annotations for a stated 1,000 scenes, and the paper does not explain how five independent annotators relate to this total (e.g., number of scenes annotated per annotator, agreement rates, or handling of multiple annotations per scene). The paper also provides no example annotation rows, no annotation guideline text, and no inter-annotator agreement metric. Without this information, readers cannot assess the stability or quality of the labels, which is load-bearing for a dataset paper.
  3. [V] The paper's own limitation statement in Section V that the retroactive instruction is 'only a proxy for a true signal' is candid, but it means the central premise -- that the instruction can serve as the cause of the observed motion -- is an unverified assumption. The authors should either validate the taxi-test heuristic (e.g., by comparing annotator-recovered instructions with actual human instructions in a small naturalistic study, or by showing that the annotated instructions predict the recorded trajectory better than a baseline) or explicitly scope the dataset as containing plausible instructions rather than causal instructions.
  4. [III] No data file, schema description, or sample of the annotation table is included in the manuscript, making it impossible to verify the format of the instructions, the handling of blank instruction fields, or the exact content of the referentiality tags. The GitHub link alone is insufficient for review; a few example rows should be included in the paper or an accessible appendix.
minor comments (5)
  1. [Abstract] In the abstract, 'A Vs' should be 'AVs' or 'autonomous vehicles.'
  2. [III] In Section III, 'thestatic reference tag' is missing a space; it should read 'the static reference tag.'
  3. [II-B] The citation formatting is inconsistent for references [8], [9], and [14]; for example, nuScenes-MQA and DRAMA use different styles for venue and year.
  4. [IV] Section IV mentions 'the first t seconds' but does not suggest a method for choosing t; a brief discussion of the temporal validity of instructions would be helpful.
  5. [III] Figure 2 is referenced without axis labels or a caption; adding a concise caption and labeled axes would improve readability.

Circularity Check

1 steps flagged · score 6.0 of 10

The central instruction–motion link in doScenes is partly circular: instructions are retroactively generated by the taxi test to match the observed trajectory, so the claimed language-to-motion connection is built into the annotation definition.

  1. self definitional [Section III, annotation protocol and taxi test; echoed in Section II-C; limitation acknowledged in Section V]
    "We apply retroactive annotation of driving instructions by playing back each of the 1,000 12-second nuScenes clips, and transcribing an instruction (or lack of instruction) that would be given to a driver from the vantage of the passenger to initiate the motion plan observed in the clip. ... if you were being driven through this scene by a taxi driver, what instruction, if any, would you need to give an instruction to trigger the behavior observed in the video?"

    The instruction annotation X is defined as an utterance that would produce the observed trajectory Y: the taxi test asks annotators to reverse-engineer a command that 'trigger[s] the behavior observed in the video.' Thus the paper's headline claim, 'creating a link between imperative language and motion,' is assured by the annotation rule rather than evidenced by independent instruction–response pairs. A model trained on doScenes sees instructions that were selected to match their paired trajectories, so the instruction-to-motion association is self-consistent by construction and cannot validate a causal or predictive language-to-driving mapping.

full rationale

The only substantive circularity is the retroactive annotation protocol: the paper's taxi-test heuristic defines an instruction as one that would initiate the observed motion in the clip, so the instruction–trajectory alignment in doScenes is guaranteed by the labeler's task. This is not hidden—Section V explicitly labels it 'a proxy for a true signal'—but it means the paper's statement that doScenes 'creates a link between imperative language and motion' is, to a substantial degree, true by construction. The paper contains no equations, fitted parameters, or evaluated predictions, so there is no fitted-input-called-prediction issue in the usual sense. The self-citations (Refs. 15-18, 27) are background and example references, not load-bearing evidence, so they do not raise the score. The absence of Talk2Car from the related-work comparison is a serious novelty/completeness concern—Talk2Car already provides real-world natural-language driving commands with referred-object annotation on nuScenes—but a missing citation or overbroad 'first' claim is a correctness risk, not a circularity. Overall score reflects partial circularity in the central dataset premise, tempered by the authors' explicit acknowledgment and by the independent referentiality-tag and instruction-diversity content.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central assumption is that retroactive instructions capture true driver-command causality. The paper itself flags this as a proxy in Section V. Additional assumptions are inherited from nuScenes: perception labels are correct, and short-term instructions remain within the egocentric viewable horizon. No free parameters or invented entities are involved because the contribution is an annotation dataset.

assumptions (3)
  • domain assumption The 'taxi test' heuristic can recover the instruction that caused the observed motion from video playback.
    Section III introduces the taxi test; Section V concedes the result is 'only a proxy for a true signal.' The dataset's core value depends on this being approximately true.
  • domain assumption nuScenes provides accurate ground truth for sensor data, 3D bounding boxes, and maps.
    Section III builds doScenes on nuScenes clips and annotations without revalidating the underlying perception labels.
  • domain assumption Relevant instruction information is contained within the egocentric viewable proximity for the defined short-term interactions.
    Section I defines short-term interactions this way, and Section IV notes that 12-second nuScenes clips can exceed the instruction's relevance.

how reviews work

0 comments
Cite this review

Pith. "Pith review of doScenes: An Autonomous Driving Dataset with Natural Language Instruction for Human Interaction and Vision-Language Navigation." pith.science (2026). https://pith.science/paper/GMXI7UM2

@misc{pith2026241205893,
  author       = {Pith},
  title        = {Pith review of: doScenes: An Autonomous Driving Dataset with Natural Language Instruction for Human Interaction and Vision-Language Navigation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GMXI7UM2}},
  note         = {Machine review of arXiv:2412.05893}
}
read the original abstract

Human-interactive robotic systems, particularly autonomous vehicles (AVs), must effectively integrate human instructions into their motion planning. This paper introduces doScenes, a novel dataset designed to facilitate research on human-vehicle instruction interactions, focusing on short-term directives that directly influence vehicle motion. By annotating multimodal sensor data with natural language instructions and referentiality tags, doScenes bridges the gap between instruction and driving response, enabling context-aware and adaptive planning. Unlike existing datasets that focus on ranking or scene-level reasoning, doScenes emphasizes actionable directives tied to static and dynamic scene objects. This framework addresses limitations in prior research, such as reliance on simulated data or predefined action sets, by supporting nuanced and flexible responses in real-world scenarios. This work lays the foundation for developing learning strategies that seamlessly integrate human instructions into autonomous systems, advancing safe and effective human-vehicle collaboration for vision-language navigation. We make our data publicly available at https://www.github.com/rossgreer/doScenes

Figures

Figures reproduced from arXiv: 2412.05893 by the authors.

Figure 1
Figure 1. Typical nuScenes data includes 3D bounding box annotations, LiDAR [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Histogram of number of instruction annotations per scene; most [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Automated Data Curation Using GPS & NLP to Generate Instruction-Action Pairs for Autonomous Vehicle Vision-Language Navigation Datasets

    cs.RO 2025-05 conditional novelty 5.0 of 10

    GPS voice directions can be automatically transcribed and synchronized with video and GPS trajectory to produce instruction-action triads for vision-language navigation training.

  2. Generative AI for Autonomous Driving: Frontiers and Opportunities

    cs.CV 2025-05 accept novelty 2.0 of 10

    A comprehensive, structured survey of generative AI for autonomous driving, covering model families, sensor modalities, real-world applications, and open research challenges.

Reference graph

Works this paper leans on

28 extracted references · 20 canonical work pages · cited by 2 Pith papers

  1. [1]

    Looking-in and looking-out of a vehicle: Computer-vision-based enhanced vehicle safety,

    M. M. Trivedi, T. Gandhi, and J. McCall, “Looking-in and looking-out of a vehicle: Computer-vision-based enhanced vehicle safety,” IEEE Trans- actions on Intelligent Transportation Systems, vol. 8, no. 1, pp. 108–120, 2007

  2. [2]

    Causal diagrams for empirical research,

    J. Pearl, “Causal diagrams for empirical research,” Biometrika, vol. 82, no. 4, pp. 669–688, 1995

  3. [3]

    Natsgd: A dataset with speech, gestures, and demonstrations for robot learning in natural human-robot interaction,

    S. Shrestha, Y . Zha, S. Banagiri, G. Gao, Y . Aloimonos, and C. Fer- muller, “Natsgd: A dataset with speech, gestures, and demonstrations for robot learning in natural human-robot interaction,” 2024

  4. [4]

    Bridgedata v2: A dataset for robot learning at scale,

    H. Walke, K. Black, A. Lee, M. J. Kim, M. Du, C. Zheng, T. Zhao, P. Hansen-Estruch, Q. Vuong, A. He, V . Myers, K. Fang, C. Finn, and S. Levine, “Bridgedata v2: A dataset for robot learning at scale,” 2024

  5. [5]

    Handmethat: Human-robot communication in physical and social environments,

    Y . Wan, J. Mao, and J. B. Tenenbaum, “Handmethat: Human-robot communication in physical and social environments,” 2023

  6. [6]

    nuscenes: A multimodal dataset for autonomous driving,

    H. Caesar, V . Bankiti, A. H. Lang, S. V ora, V . E. Liong, Q. Xu, A. Kr- ishnan, Y . Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” 2020

  7. [7]

    Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario,

    T. Qian, J. Chen, L. Zhuo, Y . Jiao, and Y .-G. Jiang, “Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, pp. 4542–4550, 2024

  8. [8]

    Nuscenes-mqa: Integrated evaluation of captions and qa for autonomous driving datasets using markup annotations,

    Y . Inoue, Y . Yada, K. Tanahashi, and Y . Yamaguchi, “Nuscenes-mqa: Integrated evaluation of captions and qa for autonomous driving datasets using markup annotations,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pp. 930–938, 2024

Show all 28 references
  1. [9]

    Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning,

    E. Sachdeva, N. Agarwal, S. Chundi, S. Roelofs, J. Li, M. Kochenderfer, C. Choi, and B. Dariush, “Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pp. 7513...

  2. [10]

    Gpt-driver: Learning to drive with gpt,

    J. Mao, Y . Qian, J. Ye, H. Zhao, and Y . Wang, “Gpt-driver: Learning to drive with gpt,” arXiv preprint arXiv:2310.01415 , 2023

  3. [11]

    Drivemlm: Aligning multi-modal large language models with behavioral planning states for autonomous driving,

    W. Wang, J. Xie, C. Hu, H. Zou, J. Fan, W. Tong, Y . Wen, S. Wu, H. Deng, Z. Li, et al., “Drivemlm: Aligning multi-modal large language models with behavioral planning states for autonomous driving,” arXiv preprint arXiv:2312.09245, 2023

  4. [12]

    Lmdrive: Closed-loop end-to-end driving with large language models,

    H. Shao, Y . Hu, L. Wang, G. Song, S. L. Waslander, Y . Liu, and H. Li, “Lmdrive: Closed-loop end-to-end driving with large language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15120–15130, 2024

  5. [13]

    Drivegpt4: Interpretable end-to-end autonomous driving via large language model,

    Z. Xu, Y . Zhang, E. Xie, Z. Zhao, Y . Guo, K.-Y . K. Wong, Z. Li, and H. Zhao, “Drivegpt4: Interpretable end-to-end autonomous driving via large language model,” IEEE Robotics and Automation Letters , 2024

  6. [14]

    Drama: Joint risk localization and captioning in driving,

    S. Malla, C. Choi, I. Dwivedi, J. H. Cho, and J. Li, “Drama: Joint risk localization and captioning in driving,” 2023

  7. [15]

    Towards explainable, safe autonomous driving with language embeddings for novelty identification and active learning: Framework and experimental analysis with real-world data sets,

    R. Greer and M. Trivedi, “Towards explainable, safe autonomous driving with language embeddings for novelty identification and active learning: Framework and experimental analysis with real-world data sets,” arXiv preprint arXiv:2402.07320, 2024

  8. [16]

    Autonomous vehicles that alert humans to take-over controls: Modeling with real-world data,

    A. Rangesh, N. Deo, R. Greer, P. Gunaratne, and M. M. Trivedi, “Autonomous vehicles that alert humans to take-over controls: Modeling with real-world data,” in2021 IEEE International Intelligent Transporta- tion Systems Conference (ITSC) , pp. 231–236, IEEE, 2021. 6

  9. [17]

    Safe control transitions: Machine vision based observable readiness index and data-driven takeover time prediction,

    R. Greer, N. Deo, A. Rangesh, M. Trivedi, and P. Gunaratne, “Safe control transitions: Machine vision based observable readiness index and data-driven takeover time prediction,” in 27th International Technical Conference on the Enhanced Safety of Vehicles (ESV) National Highwa...

  10. [18]

    Predicting take-over time for autonomous driving with real-world data: Robust data augmentation, models, and evaluation,

    A. Rangesh, N. Deo, R. Greer, P. Gunaratne, and M. M. Trivedi, “Predicting take-over time for autonomous driving with real-world data: Robust data augmentation, models, and evaluation,” arXiv preprint arXiv:2107.12932, 2021

  11. [19]

    Navila: Legged robot vision-language-action model for navigation,

    A.-C. Cheng, Y . Ji, Z. Yang, X. Zou, J. Kautz, E. Bıyık, H. Yin, S. Liu, and X. Wang, “Navila: Legged robot vision-language-action model for navigation,” arXiv preprint arXiv:2412.04453 , 2024

  12. [20]

    Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,

    P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. S ¨underhauf, I. Reid, S. Gould, and A. Van Den Hengel, “Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,” in Proceedings of the IEEE conference on computer visio...

  13. [21]

    Towards learning a generalist model for embodied navigation,

    D. Zheng, S. Huang, L. Zhao, Y . Zhong, and L. Wang, “Towards learning a generalist model for embodied navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 13624–13634, 2024

  14. [22]

    Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action,

    D. Shah, B. Osi ´nski, S. Levine, et al., “Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action,” in Conference on robot learning , pp. 492–504, PMLR, 2023

  15. [23]

    Grounding language to natural human-robot interaction in robot navigation tasks,

    Q. Xu, Y . Hong, Y . Zhang, W. Chi, and L. Sun, “Grounding language to natural human-robot interaction in robot navigation tasks,” in 2021 IEEE International Conference on Robotics and Biomimetics (ROBIO) , pp. 352–357, IEEE, 2021

  16. [24]

    Safe navigation with human instructions in complex scenes,

    Z. Hu, J. Pan, T. Fan, R. Yang, and D. Manocha, “Safe navigation with human instructions in complex scenes,” IEEE Robotics and Automation Letters, vol. 4, no. 2, pp. 753–760, 2019

  17. [25]

    Multimodal trajectory prediction conditioned on lane-graph traversals,

    N. Deo, E. Wolff, and O. Beijbom, “Multimodal trajectory prediction conditioned on lane-graph traversals,” in Conference on Robot Learning, pp. 203–212, PMLR, 2022

  18. [26]

    Thomas: Trajectory heatmap output with learned multi-agent sam- pling,

    T. Gilles, S. Sabatini, D. Tsishkou, B. Stanciulescu, and F. Moutarde, “Thomas: Trajectory heatmap output with learned multi-agent sam- pling,” in International Conference on Learning Representations , 2022

  19. [27]

    Trajectory prediction in autonomous driving with a lane heading auxiliary loss,

    R. Greer, N. Deo, and M. Trivedi, “Trajectory prediction in autonomous driving with a lane heading auxiliary loss,” IEEE Robotics and Automa- tion Letters, vol. 6, no. 3, pp. 4907–4914, 2021

  20. [28]

    Spatialrgpt: Grounded spatial reasoning in vision-language models,

    A.-C. Cheng, H. Yin, Y . Fu, Q. Guo, R. Yang, J. Kautz, X. Wang, and S. Liu, “Spatialrgpt: Grounded spatial reasoning in vision-language models,” in NeurIPS, 2024

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.