REVIEW 3 major objections 5 minor 63 references
UAV-DualCog is a benchmark that forces multimodal language models to reason about a drone's own state and its environment, grounding answers in bounding boxes and time intervals, and shows current models are far from reliable.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 21:04 UTC pith:IJJYYFH6
load-bearing objection A well-built UAV benchmark that makes self- and environment-state reasoning jointly testable and requires grounding outputs; the unvalidated simulator GT is a genuine soft spot, but the core finding that MLLMs lag humans is robust. the 3 major comments →
Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that UAV intelligence requires a coupled dual-cognition capability—self-state cognition and environment-state cognition—and that existing benchmarks do not measure this coupling. To test it, the authors build an automated pipeline that converts simulator scenes into semantic point clouds, samples UAV viewpoints and trajectories, and generates thousands of image and video question-answer pairs requiring explicit spatial or temporal grounding. On this benchmark, current MLLMs exhibit a consistent recognition–grounding gap: they can select reasonable multiple-choice answers but fail to localize landmarks or time intervals reliably. The paper further claims that the
What carries the argument
The dual-cognition formulation C(o)={c_s, c_e}, which separates self-state and environment-state reasoning, and the requirement of evidence grounding—every answer must be paired with a normalized bounding box (image tasks) or temporal interval (video tasks). The automated construction pipeline, which turns simulator scenes into scene-level semantic point clouds and samples coverage-aware viewpoints and behavior-validated flight trajectories, is the mechanism that lets the benchmark scale to thousands of annotated samples with verifiable ground truth.
Load-bearing premise
The benchmark's findings rest on the assumption that the automatically generated simulator scenes, landmark annotations, visibility intervals, and flight trajectories are accurate and representative of real UAV perception; if the simulation is erroneous or unrepresentative, the low model scores could be artifacts of the simulator rather than facts about UAV cognition. The authors explicitly note that reducing the sim-to-real gap remains future work.
What would settle it
Run the same six task types on real drone footage with accurate pose, visibility, and ground-truth annotations. If models perform substantially better or humans perform much worse on real footage than on the simulator benchmark, the claimed capability gap is at least partly a simulator artifact. Alternatively, a single well-localized case where a model correctly predicts the option but the bounding box has near-zero IoU with the landmark would illustrate the recognition–grounding gap.
If this is right
- If correct, benchmarking UAV MLLMs should always include both self- and environment-state reasoning, because the two are coupled in aerial perception.
- The recognition–grounding gap implies that MCQ accuracy overstates model competence; requiring structured evidence is a cheap way to expose true capability.
- Training on UAV-DualCog-Train improved answer accuracy by 20.8 points and bounding-box accuracy by 51.8 points on a small model, indicating synthetic evidence-grounding data can transfer to better grounded outputs.
- Humans perform far above all tested models, so the low scores reflect model limitation rather than unanswerable or ill-posed tasks.
- Current models' uneven video performance suggests temporal interval localization is a persistent bottleneck distinct from action recognition.
Where Pith is reading between the lines
- The dual-cognition separation could transfer to other embodied settings (ground robots, vehicles) where an agent must jointly track its own pose and the state of the world; a similar evidence-grounding protocol might expose analogous gaps.
- Because the pipeline is simulator-based, a natural test is whether the same gaps appear on real UAV footage with accurate pose and visibility metadata; the authors themselves defer this sim-to-real check to future work.
- The success of the optimization probe hints that reasoning about spatial evidence (bounding boxes) may partially transfer to temporal reasoning; if confirmed by further experiments, it would suggest that grounding-oriented objectives, rather than longer textual thinking, drive progress.
- The error-family analysis (dual-coordinate reasoning being the largest image-side error) suggests that the core difficulty is viewpoint transformation—models struggle to map a landmark-centric reference to an ego-centric one—which could be probed by simple 2D rotation tasks in isolation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. UAV-DualCog is a benchmark for evaluating multimodal LLMs on UAV spatio-temporal reasoning. It operationalizes 'dual cognition' as self-state vs environment-state reasoning and includes four image tasks and two video tasks, each requiring discrete answers plus normalized bounding boxes or temporal intervals. The data are built automatically from AerialVLN scenes via scene-level semantic point clouds, pose sampling, visibility checks, and manual landmark verification, yielding 4,096 image and 2,048 video samples. The paper evaluates a large set of proprietary, open-weight, and spatially specialized MLLMs, reports low grounding scores relative to answer accuracy, validates with thinking/frontier models and a human baseline, and shows a two-stage optimization probe on a disjoint-scene training split improves both answer accuracy and grounding. The central claim is that current MLLMs remain far from reliable in UAV dual cognition.
Significance. If the automatic annotations are correct, the paper makes a useful contribution: it provides a scalable benchmark that explicitly requires evidence grounding, with a broad model evaluation, a human baseline, open-ended validation, and a disjoint-scene training split. The dual-cognition taxonomy is a reasonable and timely framing, and the training probe suggests the benchmark can double as a data resource. The main strength is the breadth of diagnostics; the main weakness is the absence of quantitative validation of the automatically generated bounding boxes and temporal intervals, which are exactly the evidence outputs used to conclude that models are unreliable. If the authors add an annotation-quality study, the benchmark would be significantly stronger.
major comments (3)
- [§3.3, 'The Construction Pipeline'; §4.2, Table 5] The central conclusion that MLLMs are 'far from reliable' depends on the automatically generated normalized bounding boxes (image) and visible temporal intervals (video). The paper describes manual verification and web-based inspection but reports no quantitative validation of these GT annotations. The human baseline itself only reaches 47.1% image mIoU (Table 5), which indicates the target boxes are not trivially unambiguous; if the simulator's boxes or intervals are systematically biased, the low model grounding scores and the answer-vs-grounding gap could be artifacts. Please add an annotation-quality audit: independent human re-annotation of a random sample with IoU/tIoU agreement, re-projection errors for point-cloud-derived boxes, and per-task ambiguity statistics. This is not the deferred sim-to-real gap; it is in-simulator GT correctness, and it also affects the interpretation of
- [§4.1, 'Parsing Audit'; Table 3] The official leaderboard's strict parsing protocol makes near-zero grounding scores partly a format-compliance problem. Corrected scores are dramatically higher for the representative models (e.g., Gemini 3.1 Flash Lite mIoU 1.1→32.6, Gemini 3 Flash 1.0→23.6, GLM-4.6V 7.3→30.0). Because only 'representative low-grounding cases' are corrected, Table 2 does not tell the reader which models fail due to perception vs. JSON compliance. The qualitative conclusion may survive, but the magnitude of the grounding gap and the model rankings on localization are confounded. Please report corrected metrics for all models or a random sample, or at least the per-model format-error rate, and discuss the implications for the 'answer-vs-grounding' gap.
- [§4.3, Table 5 'Human Baseline'] No details are given on the number of human annotators, their expertise, the interface, whether the same JSON output format was required, or how disagreements were adjudicated. Since this baseline is used to argue that low model scores are not caused by invalid task design, the result is not reproducible and could be inflated or deflated by format effects. Please report the human protocol and per-task agreement, even if briefly or in the supplementary material.
minor comments (5)
- [Eq. (2)] The reward weights w_f, w_a, w_b in the joint GRPO reward are not reported. Specify the values or the search range, and ideally a small sensitivity analysis, so the optimization probe is reproducible.
- [§4.3, 'Open-ended Validation'] The 'Stage-3' and 'Stage-4' open-ended evaluations are invoked without defining what these stages are, which tasks they include, or the sample sizes. Please define them or move the details to the supplementary material.
- [Intro, §4.1] The paper calls the evaluation 'lightweight MLLMs under the computational constraints of UAV edge deployment' but evaluates frontier proprietary models such as GPT-5.5 and Gemini 3.1 Pro. Clarify the selection criterion and how 'lightweight' is meant.
- [Figure 1] Minor typos should be corrected: 'foward' → 'forward' and 'Dront' → 'Front'.
- [Table 2] The column headers 'm sIoU@50 m sIoU' are visually confusing because the same abbreviation appears in adjacent columns. Rename for clarity, e.g., 'BBox@0.5' and 'mIoU' consistently.
Circularity Check
No significant circularity: benchmark construction and evaluations are self-contained with external anchors.
full rationale
The paper does not present a derivation chain in which an output is defined or fitted from the very quantity it claims to predict. The central claims are empirical statements about MLLM performance on a constructed benchmark. Ground-truth bounding boxes and temporal intervals are generated from the AerialVLN simulator via pose sampling, semantic point-cloud fusion, visibility checks, and manual verification; whether these annotations are accurate is a data-quality concern, not an input-output identity. The human baseline (image 80.4% acc / 47.1% mIoU; video 60.8% acc / 53.1% mtIoU) and the disjoint train/test split provide independent external anchors. The optimization probe trains on six unused scenes and evaluates on the held-out test split, so the reported gains (e.g., +20.78 pp option accuracy, +51.76 pp BBox@0.5) are held-out measurements rather than forced by construction. No load-bearing self-citation appears: the cited AerialVLN and related benchmarks are external resources, and the concurrent SIS-Bench [63] is not used to justify this paper's results. The deferred sim-to-real gap and the lack of quantitative annotation validation are correctness risks, not circularity. No specific step exhibits the required reduction of a claimed result to its own inputs.
Axiom & Free-Parameter Ledger
free parameters (1)
- Reward weights w_f, w_a, w_b in Eq. (2) =
not reported
axioms (5)
- domain assumption Flight trajectories and observations generated in AerialVLN scenes are sufficiently representative of real UAV multiview/spatio-temporal conditions.
- domain assumption Automated semantic point-cloud fusion, visibility checking, and manual landmark verification yield correct ground-truth poses, bounding boxes, and intervals.
- ad hoc to paper The dual-cognition taxonomy (self-state vs environment-state) and the six task types are a valid operationalization of UAV reasoning.
- domain assumption BBox IoU>=0.5 and tIoU>=0.5 are appropriate evidence measures, and low grounding scores are not primarily artifacts of strict JSON parsing.
- domain assumption The human baseline annotations are accurate and representative.
read the original abstract
Multimodal large language models have achieved strong performance across diverse vision-language tasks, yet their capabilities in UAV scenarios remain insufficiently explored. Recent UAV-oriented benchmarks have begun to evaluate MLLMs in aerial scenarios, but they typically focus on scene understanding, event recognition, or navigation completion, rather than jointly assessing the dual-cognition capability required for UAV agents: reasoning about both the UAV's own state and the external environment in multiview spatio-temporal contexts. To address this gap, we present UAV-DualCog, a benchmark for aerial multiview spatio-temporal reasoning built on this dual-cognition perspective. UAV-DualCog includes both image and video tasks to jointly evaluate self-state and environment-state reasoning, while requiring spatial or temporal grounding beyond discrete answer prediction. We also develop an automated pipeline that constructs data from scene-level semantic point clouds, yielding a scalable benchmark with diverse scenes, hundreds of landmarks, and thousands of QA samples. Extensive evaluations show that current MLLMs remain far from reliable in UAV dual cognition. Self-state reasoning, viewpoint transformation, precise spatial grounding, and temporal interval localization are persistent bottlenecks, and additional validation with thinking/frontier models and a human baseline confirms that the benchmark is understandable to humans but challenging for existing models. We further construct UAV-DualCog-Train from disjoint scenes and show through a lightweight optimization probe that it provides useful structured supervision, suggesting its value not only as an evaluation benchmark but also as a data resource for advancing MLLM-based UAV agents. Project website and supplementary materials: https://uav-dualcog.lozumi.com
Figures
Reference graph
Works this paper leans on
-
[1]
Remyx AI and Salma Mayorquin. 2025. SpaceOm Models. https://huggingface. co/remyxai/SpaceOm
2025
-
[2]
Remyx AI and Salma Mayorquin. 2025. SpaceThinker Models. https:// huggingface.co/remyxai/SpaceThinker-Qwen2.5VL-3B
2025
-
[3]
Anthropic. 2026. Introducing Claude Opus 4.6. https://www.anthropic.com/ news/claude-opus-4-6
2026
-
[4]
Anthropic. 2026. Introducing Claude Opus 4.7. https://www.anthropic.com/ news/claude-opus-4-7
2026
-
[5]
Anthropic. 2026. Introducing Claude Opus 4.8. https://www.anthropic.com/ news/claude-opus-4-8
2026
-
[6]
Anthropic. 2026. Introducing Claude Sonnet 4.6. https://www.anthropic.com/ news/claude-sonnet-4-6
2026
-
[7]
Mohammadamin Barekatain, Miquel Martí, Hsueh-Fu Shih, Samuel Murray, Kotaro Nakayama, Yutaka Matsuo, and Helmut Prendinger. 2017. Okutama- action: An aerial view video dataset for concurrent human action detection. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops. 28–35
2017
-
[8]
Zhongang Cai, Ruisi Wang, Chenyang Gu, Fanyi Pu, Junxiang Xu, Yubo Wang, Wanqi Yin, Zhitao Yang, Chen Wei, Qingping Sun, Tongxi Zhou, Jiaqi Li, Hui En Pang, Oscar Qian, Yukun Wei, Zhiqian Lin, Xuanke Shi, Kewang Deng, Xiaoyang Han, Zukai Chen, Xiangyu Fan, Hanming Deng, Lewei Lu, Liang Pan, Bo Li, Ziwei Liu, Quan Wang, Dahua Lin, and Lei Yang. 2025. Scali...
arXiv 2025
-
[9]
Google DeepMind. 2025. Gemini 3 Flash: frontier intelligence built for speed. https://blog.google/products-and-platforms/products/gemini/gemini-3-flash/
2025
-
[10]
Google DeepMind. 2026. Gemini 3.1 Flash-Lite: Built for intelligence at scale. https://blog.google/innovation-and-ai/models-and-research/gemini- models/gemini-3-1-flash-lite/
2026
-
[11]
Google DeepMind. 2026. Gemini 3.1 Pro: A smarter model for your most com- plex tasks. https://blog.google/innovation-and-ai/models-and-research/gemini- models/gemini-3-1-pro/
2026
-
[12]
Dawei Du, Yuankai Qi, Hongyang Yu, Yifan Yang, Kaiwen Duan, Guorong Li, Weigang Zhang, Qingming Huang, and Qi Tian. 2018. The unmanned aerial vehicle benchmark: Object detection and tracking. InProceedings of the European conference on computer vision (ECCV). 370–386
2018
-
[13]
Yue Fan, Winson Chen, Tongzhou Jiang, Chun Zhou, Yi Zhang, and Xin Eric Wang. 2022. Aerial Vision-and-Dialog Navigation.arXiv preprint arXiv:2205.12219 (2022)
Pith/arXiv arXiv 2022
-
[14]
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, et al. 2023. Mme: A comprehensive evaluation benchmark for multimodal large language models.arXiv preprint arXiv:2306.13394(2023)
Pith/arXiv arXiv 2023
-
[15]
Yunpeng Gao, Chenhui Li, Zhongrui You, Junli Liu, Zhen Li, Pengan Chen, Qizhi Chen, Zhonghan Tang, Liansheng Wang, Penghui Yang, Yiwen Tang, Yuhang Tang, Shuai Liang, Songyi Zhu, Ziqin Xiong, Yifei Su, Xinyi Ye, Jianan Li, Yan Ding, Dong Wang, Zhigang Wang, Bin Zhao, and Xuelong Li. 2025. OpenFly: A Comprehensive Platform for Aerial Vision-Language Naviga...
arXiv 2025
-
[16]
Wentao Ge, Shunian Chen, Hardy Chen, Nuo Chen, Junying Chen, Zhihong Chen, Wenya Xie, Shuo Yan, ChenghaoZhu ChenghaoZhu, Ziyue Lin, et al. 2025. Mllm-bench: evaluating multimodal llms with per-sample criteria. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Techno...
2025
-
[17]
Jingyu Guo, Ziye Chen, Ziwen Li, Zhengqing Gao, Jiaxin Huang, Hanlue Zhang, Fengming Huang, Yu Yao, Tongliang Liu, and Mingming Gong. 2026. HUGE- Bench: A Benchmark for High-Level UAV Vision-Language-Action Tasks.arXiv preprint arXiv:2603.19822(2026)
Pith/arXiv arXiv 2026
-
[18]
Longzhen Han, Awes Mubarak, Almas Baimagambetov, Nikolaos Polatidis, and Thar Baker. 2025. A Survey of Generative Categories and Techniques in Multi- modal Generative Models.arXiv preprint arXiv:2506.10016(2025)
arXiv 2025
-
[19]
Chih Yao Hu, Yang-Sen Lin, Yuna Lee, Chih-Hai Su, Jie-Ying Lee, Shr-Ruei Tsai, Chin-Yang Lin, Kuan-Wen Chen, Tsung-Wei Ke, and Yu-Lun Liu. 2025. See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation. InProceedings of The 9th Conference on Robot Learning (Proceedings of Machine Learning Research, Vol. 305), Joseph Lim, Shu...
2025
-
[20]
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card.arXiv preprint arXiv:2410.21276(2024)
Pith/arXiv arXiv 2024
-
[21]
Yatai Ji, Zhengqiu Zhu, Yong Zhao, Beidan Liu, Chen Gao, Yihao Zhao, Sihang Qiu, Yue Hu, and Quanjun Yin. 2026. Towards autonomous uav visual object search in city space: Benchmark and agentic methodology. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 18342–18350
2026
-
[22]
Wen Jiang, Li Wang, Kangyao Huang, Wei Fan, Jinyuan Liu, Shaoyu Liu, Hongwei Duan, Bin Xu, and Xiangyang Ji. 2025. LongFly: Long-Horizon UAV Vision-and- Language Navigation with Spatiotemporal Context Integration.arXiv preprint arXiv:2512.22010(2025)
arXiv 2025
-
[23]
Jungdae Lee, Taiki Miyanishi, Shuhei Kurita, Koya Sakamoto, Daichi Azuma, Yutaka Matsuo, and Nakamasa Inoue. 2025. Citynav: A large-scale dataset for real- world aerial navigation. InProceedings of the IEEE/CVF International Conference on Computer Vision. 5912–5922
2025
-
[24]
Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan. 2024. Seed-bench: Benchmarking multimodal large language models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13299–13308
2024
-
[25]
Tianjiao Li, Jun Liu, Wei Zhang, Yun Ni, Wenqian Wang, and Zhiheng Li. 2021. Uav-human: A large benchmark for human behavior understanding with un- manned aerial vehicles. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 16266–16275
2021
-
[26]
Yunxin Li, Zhenyu Liu, Zitao Li, Xuanyu Zhang, Zhenran Xu, Xinyu Chen, Haoyuan Shi, Shenyuan Jiang, Xintong Wang, Jifang Wang, et al. 2025. Percep- tion, reason, think, and plan: A survey on large multimodal reasoning models. arXiv preprint arXiv:2505.04921(2025)
Pith/arXiv arXiv 2025
-
[27]
Huizhi Liang, Yichao Shen, Yu Deng, Sicheng Xu, Zhiyuan Feng, Tong Zhang, Yaobo Liang, and Jiaolong Yang. 2026. HiSpatial: Taming Hierarchical 3D Spa- tial Understanding in Vision-Language Models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2026
-
[28]
Ming-Yi Lin, Ou-Wen Lee, and Chih-Ying Lu. 2024. Embodied ai with large lan- guage models: A survey and new hri framework. In2024 International Conference on Advanced Robotics and Mechatronics (ICARM). IEEE, 978–983
2024
-
[29]
Shubo Liu, Hongsheng Zhang, Yuankai Qi, Peng Wang, Yanning Zhang, and Qi Wu. 2023. AerialVLN: Vision-and-language Navigation for UAVs. InInternational Conference on Computer Vision (ICCV)
2023
-
[30]
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al . 2024. Mmbench: Is your multi-modal model an all-around player?. InEuropean conference on computer vision. Springer, 216–233
2024
-
[31]
Taiki Miyanishi, Fumiya Kitamori, Shuhei Kurita, Jungdae Lee, Motoaki Kawan- abe, and Nakamasa Inoue. 2023. CityRefer: Geography-aware 3D Visual Ground- ing Dataset on City-scale Point Cloud Data. InThirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track
2023
-
[32]
OpenAI. 2026. GPT-5.3 Instant: Smoother, more useful everyday conversations. https://openai.com/index/gpt-5-3-instant
2026
-
[33]
OpenAI. 2026. Introducing GPT-5.4. https://openai.com/index/introducing-gpt- 5-4/
2026
-
[34]
OpenAI. 2026. Introducing GPT-5.5. https://openai.com/index/introducing-gpt- 5-5
2026
-
[35]
Kun Ouyang, Yuanxin Liu, Haoning Wu, Yi Liu, Hao Zhou, Jie Zhou, Fandong Meng, and Xu Sun. 2025. SpaceR: Reinforcing MLLMs in Video Spatial Reasoning. arXiv preprint arXiv:2504.01805(2025)
Pith/arXiv arXiv 2025
-
[36]
Qwen Team. 2026. Qwen3.5: Towards Native Multimodal Agents. https://qwen. ai/blog?id=qwen3.5
2026
-
[37]
Qwen Team. 2026. Qwen3.6-Flash. https://www.qianwenai.com/models/qwen3. 6-flash
2026
-
[38]
Qwen Team. 2026. Qwen3.6-Plus: Towards Real World Agents. https://qwen.ai/ blog?id=qwen3.6
2026
-
[39]
Qwen Team. 2026. Qwen3.7-Plus: Multimodal Agent Intelligence. https://qwen. ai/blog?id=qwen3.7-plus
2026
-
[40]
Kimi Team. 2026. Kimi K2.6: Advancing Open-Source Coding. https://www. kimi.com/blog/kimi-k2-6
2026
-
[41]
Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, SH Cai, Yuan Cao, Y Charles, HS Che, Cheng Chen, Guanduo Chen, et al. 2026. Kimi K2. 5: Visual Agentic Intelligence.arXiv preprint arXiv:2602.02276(2026)
Pith/arXiv arXiv 2026
-
[42]
V Team, Wenyi Hong, Wenmeng Yu, Xiaotao Gu, Guo Wang, Guobing Gan, Haomiao Tang, Jiale Cheng, Ji Qi, Junhui Ji, Lihang Pan, Shuaiqi Duan, Weihan Wang, Yan Wang, Yean Cheng, Zehai He, Zhe Su, Zhen Yang, Ziyang Pan, Aohan Zeng, Baoxu Wang, Bin Chen, Boyan Shi, Changyu Pang, Chenhui Zhang, Da Yin, Fan Yang, Guoqing Chen, Jiazheng Xu, Jiale Zhu, Jiali Chen, J...
Pith/arXiv arXiv 2025
-
[43]
Xiaomi Mimo Team. 2026. Xiaomi MiMo-V2-Omni. https://mimo.xiaomi.com/ mimo-v2-omni
2026
-
[44]
Xiaomi Mimo Team. 2026. Xiaomi MiMo-V2.5. https://mimo.xiaomi.com/mimo- v2-5
2026
-
[45]
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. 2024. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution.arXiv preprint arXiv:2409.12191(2024)
Pith/arXiv arXiv 2024
-
[46]
Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, et al . 2025. InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Effi- ciency.arXiv preprint arXiv:2508.18265(2025)
Pith/arXiv arXiv 2025
-
[47]
Junfei Wu, Jian Guan, Kaituo Feng, Qiang Liu, Shu Wu, Liang Wang, Wei Wu, and Tieniu Tan. 2025. Reinforcing spatial reasoning in vision-language models with interwoven thinking and visual drawing.arXiv preprint arXiv:2506.09965 (2025)
Pith/arXiv arXiv 2025
-
[48]
Ruipu Wu, Yige Zhang, Jinyu Chen, Linjiang Huang, Shifeng Zhang, Xu Zhou, Liang Wang, and Si Liu. 2025. AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation. InProceedings of the 33rd ACM International Conference on Multimedia. 2576–2585
2025
-
[49]
xAI. 2025. Grok 4.1 Fast and Agent Tools API. https://x.ai/news/grok-4-1-fast
2025
-
[50]
Huilin Xu, Zhuoyang Liu, Yixiang Luomei, and Feng Xu. 2025. Aerial Vision- Language Navigation with a Unified Framework for Spatial, Temporal and Em- bodied Reasoning.arXiv preprint arXiv:2512.08639(2025)
Pith/arXiv arXiv 2025
-
[51]
Jihan Yang, Shusheng Yang, Anjali W Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie. 2025. Thinking in space: How multimodal large language models see, remember, and recall spaces. InProceedings of the Computer Vision and Pattern Recognition Conference. 10632–10643
2025
-
[52]
Rui Yang, Ziyu Zhu, Yanwei Li, Jingjia Huang, Shen Yan, Siyuan Zhou, Zhe Liu, Xiangtai Li, Shuangye Li, Wenqian Wang, et al. 2025. Visual spatial tuning.arXiv preprint arXiv:2511.05491(2025)
arXiv 2025
-
[53]
Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. 2023. Mm-vet: Evaluating large multimodal models for integrated capabilities.arXiv preprint arXiv:2308.02490(2023)
Pith/arXiv arXiv 2023
-
[54]
Weihao Yu, Zhengyuan Yang, Lingfeng Ren, Linjie Li, Jianfeng Wang, Kevin Lin, Chung-Ching Lin, Zicheng Liu, Lijuan Wang, and Xinchao Wang. 2024. Mm-vet v2: A challenging benchmark to evaluate large multimodal models for integrated capabilities.arXiv preprint arXiv:2408.00765(2024)
Pith/arXiv arXiv 2024
-
[55]
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. 2024. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 9556–9567
2024
-
[56]
Fanlong Zeng, Wensheng Gan, Yongheng Wang, Ning Liu, and Philip S Yu. 2023. Large language models for robotics: A survey.arXiv preprint arXiv:2311.07226 (2023)
arXiv 2023
-
[57]
Weichen Zhang, Chen Gao, Shiquan Yu, Ruiying Peng, Baining Zhao, Qian Zhang, Jinqiang Cui, Xinlei Chen, and Yong Li. 2025. Citynavagent: Aerial vision-and- language navigation with hierarchical semantic planning and global memory. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 31292–31309
2025
-
[58]
Weichen Zhang, Zile Zhou, Xin Zeng, Liu Xuchen, Jianjie Fang, Chen Gao, Jinqiang Cui, Yong Li, Xinlei Chen, and Xiao-Ping Zhang. 2025. Open3d-vqa: A benchmark for embodied spatial concept reasoning with multimodal large language model in open space. InProceedings of the 33rd ACM International Conference on Multimedia. 12784–12791
2025
-
[59]
Xinyuan Zhang, Yonglin Tian, Fei Lin, Yue Liu, Jing Ma, Xiao Wang, Kornélia Sára Szatmáry, and Fei-Yue Wang. 2025. LogisticsVLN: Vision-language navigation for low-altitude terminal delivery based on agentic UAVs. In2025 IEEE 28th International Conference on Intelligent Transportation Systems (ITSC). IEEE, 4437– 4442
2025
-
[60]
Baining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang, Jirong Zha, Weichen Zhang, Chen Gao, Yue Wang, Jinqiang Cui, Xinlei Chen, and Yong Li. 2025. UrbanVideo- Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces. arXiv:2503.06157 [cs.CV] https://arxiv.org/abs/ 2503.06157
arXiv 2025
-
[61]
Yong Zhao, Kai Xu, Zhengqiu Zhu, Yue Hu, Zhiheng Zheng, Yingfeng Chen, Yatai Ji, Chen Gao, Yong Li, and Jincai Huang. 2025. Cityeqa: A hierarchical llm agent on embodied question answering benchmark in city space. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 12476–12491
2025
-
[62]
Yicheng Zou, Dongsheng Zhu, Lin Zhu, Tong Zhu, Yunhua Zhou, Peiheng Zhou, Xinyu Zhou, Dongzhan Zhou, Zhiwang Zhou, Yuhao Zhou, Bowen Zhou, Zhanping Zhong, Zhijie Zhong, Haiteng Zhao, Penghao Zhao, Xiaomeng Zhao, Zhiyuan Zhao, Yechen Zhang, Jin Zhang, Wenwei Zhang, Hongjie Zhang, Zhuo Zhang, Wenlong Zhang, Bo Zhang, Chao Zhang, Chen Zhang, Yuhang Zang, Fei...
arXiv 2026
-
[63]
Zhishan Zou, Guoyan Sun, Zhiwei Wei, Jiancheng Pan, Yujie Li, Mugen Peng, and Wenjia Xu. 2026. Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence.arXiv preprint arXiv:2607.12477(2026). https://arxiv.org/abs/2607.12477
Pith/arXiv arXiv 2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.