Pith. sign in

REVIEW 3 major objections 4 minor 63 references

360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents

T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Best AI agent scores 17.1% where humans hit 77.3% on Akihabara navigation

desk verdict A careful benchmark with a plausible strong LMM-human gap, but the unvalidated inherited pose graph and missing artifacts make the 'real urban navigation' claim premature. read the letter →

arxiv 2608.08814 v1 pith:IOQGZVHC submitted 2026-08-09 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords embodiedAIurbannavigationbenchmark360-degreevideophotorealisticvirtualenvironmentvision-languagespatialreasoningAkihabara
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces 360CityArena, a benchmark for embodied agents navigating a photorealistic virtual reconstruction of Akihabara, Tokyo, built from 602 interconnected 360-degree video segments covering 85 streets. It contains 175 manually crafted tasks across three categories — environment understanding, path reasoning, and spatial reasoning — and reports that the strongest evaluated multimodal agent, Gemini 2.5 Flash, solves only 17.1% of tasks while local human participants solve 77.3%. The aim is to provide a realistic, dynamic, street-level testbed that existing outdoor simulators and Street-View-based environments lack, and to measure how far current agents are from usable urban exploration. A sympathetic reader would take the paper as showing that city-scale embodied navigation remains largely unsolved, even by large language-vision models.

What carries the argument

The central object is the Realistic Virtual World of Akihabara: 602 360-degree video segments projected onto spheres and linked into a navigable pose graph in Unity, with 193 nodes and 305 edges over 85 streets, giving agents smooth motion along filmed trajectories with directional capture per street. This pose graph is the substrate for all 175 tasks. The evaluation also rests on four metrics: exact match for grid-coordinate answers, fuzzy match with an isolated language-model judge for relational answers, coordinate match by Euclidean distance for navigation tasks, and mean relative accuracy for counting tasks. The design enables controlled comparisons, such as identical landmark-search tasks differing only in whether the goal is given as text or as an image.

What would settle it

Take a random sample of 360CityArena tasks, run them in the real Akihabara district with a human or instrumented recorder, and check whether every ground-truth answer — landmark identity, grid cell, route reachability, and object count — matches what is actually on the ground. Any systematic mismatch, such as a Map Navigation route that cuts through a building or a landmark appearing at the wrong corner, would invalidate the benchmark as a real-world proxy.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central claim is that a pose graph of 360-degree video trajectories, grounded in the real street network of Akihabara, can serve as a valid urban-district benchmark and that current state-of-the-art LMM-based agents fall far short of human-level performance on it. The claim is supported by a task suite of 175 human-verified tasks, four evaluation protocols, and experiments with six proprietary and open models. The paper also reports that image-based landmark search generally beats language-based landmark search for the same landmarks, that performance declines with task difficulty, and that giving an agent its current location on a map does not consistently help and sometimes hurts.

Load-bearing premise

The benchmark's value as a proxy for real urban navigation depends on the 360-degree video reconstruction faithfully representing Akihabara's street connectivity, landmark placements, and visual appearance; if the pose graph distorts the layout, the ground-truth answers and the human baseline would not transfer to real streets.

Editorial extensions

If this is right

  • If the benchmark is accepted, no current LMM-based agent is close to human-level urban exploration: the best model reaches 17.1% overall versus 77.3% for local human participants.
  • Image-based landmark search is generally easier than language-based landmark search for LMM agents, so visual goal specification is a more reliable channel than textual landmark descriptions.
  • Performance degrades monotonically as tasks move from Easy to Hard, so the benchmark can rank models by difficulty and expose where specific abilities break down.
  • Providing self-location as a map marker does not reliably improve agents and can degrade relational spatial reasoning, implying that map-to-view alignment is a distinct, still-unsolved capability.
  • Because all task categories use the same photorealistic environment, the benchmark supports direct comparison of perception, planning, and counting abilities under matched visual conditions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the human–agent gap holds in follow-up work, the bottleneck is less perception than exploration and grounding: the paper's failure analysis attributes many errors to action overshoot, stagnation loops, and aligning map information with first-person views rather than to object recognition alone.
  • The same pose-graph construction should transfer to other cities using the Movie Map paradigm, so a multi-city version would directly test whether the observed gap is specific to Akihabara's visual density or general to urban navigation.
  • A testable extension is to vary the step limit or the availability of the map marker to see whether the low scores reflect poor planning rather than limited observation; the current stopping conditions already suggest stagnation and wrong actions are major factors.
  • The 11.3% share of boundary-transition actions could be probed by comparing agent success on trajectories that cross video-seam boundaries versus trajectories that stay within a single filmed segment, isolating any artifact of the benchmark's video construction.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces 360CityArena, an embodied urban navigation benchmark built on a 360-degree video reconstruction of Akihabara, Tokyo, comprising 175 tasks across three categories and seven subcategories: Environment Understanding, Path Reasoning, and Spatial Reasoning. The authors evaluate six LMM-based agents (GPT-5, Claude Sonnet 4.5, Gemini 2.5 Flash, Qwen2.5-VL, and two InternVL variants) and five local-expert human participants, reporting a large performance gap (best model Gemini 2.5 Flash at 17.1% versus humans at 77.3%). The benchmark uses four evaluation metrics — exact match, fuzzy match with a validated GPT-5 judge, coordinate match with threshold sensitivity analysis, and mean relative accuracy — and includes analyses of location-information effects and failure modes. The paper also provides a detailed appendix with task prompts and experimental settings.

Significance. If the environment fidelity holds, 360CityArena is a valuable addition to embodied AI benchmarking: it combines photorealistic, dynamic, city-scale observations with seven task types covering perception, path reasoning, and spatial reasoning. The evaluation protocol is careful and transparent: the fuzzy-match judge is validated against human agreement (kappa 0.937), coordinate thresholds are checked for sensitivity (Kendall's tau >= 0.89), and task statistics monotonically support the Easy/Medium/Hard labels. The public benchmark infrastructure, detailed prompts, and reproducible metric definitions are strengths that make the paper useful regardless of the specific model results. However, the benchmark's real-world validity hinges on the fidelity of the underlying reconstructed pose graph, which is not independently validated in this manuscript.

major comments (3)
  1. [Section 3.1, Section 6, Section A.2] The validity of the benchmark as a test of real urban navigation rests on the inherited Realistic Virtual World pose graph from reference [41], but the manuscript does not validate that pose graph's topology or coordinate accuracy. The 193-node, 305-edge graph is described, and Section 6 only qualitatively checks boundary discontinuities; there is no comparison of the pose graph to the actual Akihabara street network or to the OpenStreetMap crops used in Map Navigation, Localization, and coordinate-match ground truth. Without such validation, the optimal-route labels, grid-cell answers, and coordinate thresholds (10 m for most tasks, 20 m for Map Navigation) may be miscalibrated, and the human Map Navigation result (92%) could reflect local-knowledge compensation for a distorted map rather than valid benchmark measurement. Please add a quantitative validation, such as node/edge precision and recall against OSM street segments and coordinate error at intersections, or explicitly cite where reference [41] provides this validation and summarize its results.
  2. [Section 5.1, Table 2] The human baseline consists of only five participants, and no variance, confidence interval, or per-participant breakdown is reported. Since the headline claim "human 77.3% vs. Gemini 2.5 Flash 17.1%" is a central result, the current presentation overstates its precision. Please report confidence intervals (or per-participant scores) and clarify how the 25 tasks per subcategory were allocated across participants; if the five participants each attempted all tasks, report the between-subject variance, and if tasks were split, report the per-task sample size.
  3. [Section 5.2, Table 3] The analysis of the effect of removing location information compares cells of only 25 tasks each, with no significance testing or confidence intervals. Differences such as Landmark (Language) 24.0% vs. 16.0% and Relational Reasoning 48.0% vs. 32.0% are within plausible sampling error for n=25 (approximate standard error of 8–10 percentage points). The conclusions drawn about map-to-visual alignment difficulty are therefore not supported by the data. Please add significance tests (e.g., exact binomial or bootstrap confidence intervals) or restrict the discussion to descriptive observations.
minor comments (4)
  1. [Section 1] The sentence "360CityArena would also serves as a practical guideline" contains a grammatical error; it should be "would also serve as."
  2. [Section 5.2, Figure 5] The failure-cause breakdown uses Gemini 3 Flash to label failures, with humans verifying each label, but the paper does not report the number of classified failures or the inter-annotator agreement between Gemini 3 Flash and the human verifiers; please add this information to assess the reliability of the failure analysis.
  3. [Section 3.3] The paper states that tasks are labeled Easy, Medium, or Hard, but it does not report the number of tasks in each difficulty level per subcategory; adding this distribution would help readers interpret Figure 4 and the difficulty-related statistics.
  4. [Section B.1 and Section 4.2] The action described as "reset the viewpoint to align with the current heading direction" in Section 4.2 is referred to as "S" and labeled "turn camera to the direction of travel" in the system prompt; unify the terminology to avoid confusion for readers and for future implementations.

Circularity Check

1 steps flagged · score 2.0 of 10

No prediction in 360CityArena reduces to its inputs; the only self-citation issue is that the benchmark's realism premise is inherited from the authors' own prior RVW paper [41] without independent topology validation.

  1. self citation load bearing [Section 3.1 (Akihabara Virtual Environment); also Sections 1, 2, and 6]
    "The environment used in 360CityArena is a Realistic Virtual World of Akihabara constructed from 602 360° video segments over 85 streets [41]."

    360CityArena's realism claim and all task ground truths are defined on the pose graph of [41], whose author list overlaps with this paper (Takenawa, Aizawa). The manuscript reports graph statistics (193 nodes, 305 edges) and a qualitative boundary check, but no comparison of the pose graph to OpenStreetMap topology or to real-world coordinates, so the assertion that the benchmark measures realistic urban navigation is imported from the authors' own prior reconstruction. This is a load-bearing self-citation for the environment's fidelity, although the headline LMM-versus-human results are independent external measurements and do not reduce to [41].

full rationale

I find no derivation in 360CityArena that reduces to its own inputs by construction. The benchmark is a measurement instrument: the 175 tasks were hand-crafted by annotators in Unity, ground truths are defined inside the constructed environment, and the model/human scores in Tables 2 and 3 are external observations rather than values fitted from the environment. The coordinate_match threshold epsilon is supported by a sensitivity check (Kendall's tau >= 0.89; human-best-LMM gap above 36 pp even at epsilon = 30 m), and the fuzzy_match judge is isolated from the agents and validated against human agreement (kappa = 0.937). The one structural dependency is the environment itself: Section 3.1 adopts the Realistic Virtual World of [41], whose authors overlap with the present paper, and the paper does not independently verify the pose graph's street topology or coordinate accuracy relative to OpenStreetMap. That is a self-citation/validation gap underlying the benchmark's realism claim, not a circular reduction of any prediction to its input; therefore the score is 2 rather than 0.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests mostly on the prior Realistic Virtual World environment, OpenStreetMap data, and the annotators' judgment. The only hand-chosen numerical parameters are evaluation thresholds and stopping conditions. No invented entities are introduced.

free parameters (3)
  • coordinate_match distance threshold epsilon = 10 m; 20 m for Map Navigation
    Selected by hand as the pass/fail threshold for coordinate-based tasks. A sensitivity check shows rankings are stable from 0.75x to 1.5x, but the absolute success rates and the human-model gap depend on this choice.
  • MRA threshold set C = {0.5, 0.55, ..., 0.95}
    Selected tolerance grid for the Object Count metric. MRA scores are an average over this hand-chosen set rather than a parameter-free measure.
  • Stopping conditions = 50 steps; 20 repeated actions; 5 away-from-goal movements
    Hand-set stopping rules in Appendix A.3 that bound how long agents can explore. They affect success rates on longer tasks and are not derived from data.
assumptions (4)
  • domain assumption The Realistic Virtual World from reference [41] accurately represents Akihabara's street layout and landmarks.
    Section 3.1 builds all tasks and ground truths on this reconstruction; if the pose graph or video appearance is inaccurate, task ground truths do not transfer to the real district.
  • domain assumption Pre-recorded 360-degree video trajectories are a valid substitute for continuous physical movement for the evaluated skills.
    Section 3.1 states agents navigate along the pose graph and cannot move freely or interact; Section 6 acknowledges boundary discontinuities. The benchmark's construct validity assumes this limitation does not undermine navigation and reasoning measurements.
  • domain assumption OpenStreetMap data correctly represents road connectivity for map-based tasks.
    Appendix A.2 says maps are cropped from OpenStreetMap; Localization and Map Navigation ground truth depend on this map being correct.
  • domain assumption Annotators who had visited Akihabara can reliably create solvable tasks with correct answers.
    Section 3.3 reports eight annotators, two per subtask, one author on all tasks, and about 60 hours of construction; if annotator bias or ambiguity persists, difficulty labels and ground truth may be unreliable.

how reviews work

0 comments
Cite this review

Pith. "Pith review of 360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents." pith.science (2026). https://pith.science/paper/IOQGZVHC

@misc{pith2026260808814,
  author       = {Pith},
  title        = {Pith review of: 360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IOQGZVHC}},
  note         = {Machine review of arXiv:2608.08814}
}
read the original abstract

We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Existing outdoor benchmarks either lack sufficient photorealism or complexity, resulting in a considerable gap from real-world urban environments. 360CityArena is built on a realistic reconstruction of the Akihabara district in Tokyo, Japan, using 602 360-degree video segments covering 85 streets, and consists of 175 meticulously human-crafted tasks. It encompasses three task categories: Environment Understanding, Path Reasoning, and Spatial Reasoning, covering fundamental abilities required for urban exploration, such as localization, landmark search, path planning, and relational spatial reasoning, thereby enabling comprehensive evaluation in realistic urban scenes. Our evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level (human: 77.3% vs. Gemini 2.5 Flash: 17.1%), revealing substantial challenges that remain in city-scale embodied navigation and reasoning. 360CityArena provides a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning.

Figures

Figures reproduced from arXiv: 2608.08814 by the authors.

Figure 1
Figure 1. 360CityArena. We introduce a benchmark for evaluating embodied agents in a photorealistic reconstruction of Akihabara, Tokyo, Japan, built from interconnected 360° video trajectories. The benchmark covers realistic urban streets and evaluates agents on diverse tasks requiring environment understanding, path reasoning, and spa￾tial reasoning. Abstract. We present 360CityArena, a benchmark for evaluating the urban exp… view at source ↗
Figure 2
Figure 2. Examples in each task type in 360CityArena. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Example of the visual observations in 360CityArena. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Success rate by difficulty level across tasks and models (%). [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: Failure cause breakdown by task category across models (%). [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: Example of the agent’s views for Landmark Search with Language, resulting in failure. The agent is instructed to search for “Jonathan”. From t = 8 to t = 38, the gaze repeatedly moved up and down. At t = 39, the gaze briefly shifted to the right, but from t = 41 onward…
Figure 7
Figure 7. Figure 7: Example of the agent’s views for Landmark Search with Image, re￾sulting in success. The agent is instructed to search for “Jonathan” given in image form. 7 Conclusions We present 360CityArena, a benchmark designed to evaluate the urban ex￾ploration capabilities of embo…
Figure 8
Figure 8. Figure 8: Grid map provided in the Localization Task. [PITH_FULL_IMAGE:figures/full_fig_p024_8.png]
Figure 9
Figure 9. Figure 9: Example of a landmark image provided in the Landmark Search [PITH_FULL_IMAGE:figures/full_fig_p026_9.png]
Figure 10
Figure 10. Figure 10: Example of a map provided in the Map Navigation Task. [PITH_FULL_IMAGE:figures/full_fig_p027_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

63 extracted references · 54 canonical work pages

  1. [41]

    MTAP85, 149 (2026)

    Takenawa, M., Sugimoto, N., Wöhler, L., Ikehata, S., Aizawa, K.: Building and evaluating a realistic virtual world for large scale urban exploration from 360° videos. MTAP85, 149 (2026)

  2. [1]

    arXiv preprint arXiv:1807.06757 (2018)

    Anderson, P., Chang, A., Chaplot, D.S., Dosovitskiy, A., Gupta, S., Koltun, V., Kosecka, J., Malik, J., Mottaghi, R., Savva, M., et al.: On evaluation of embodied navigation agents. arXiv preprint arXiv:1807.06757 (2018)

  3. [2]

    Anthropic: Claude Sonnet 4.5 system card. Tech. rep., Anthropic, PBC (2025)

  4. [3]

    arXiv preprint arXiv:2502.13923 (2025)

    Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H., Lin, J.: Qwen2.5-VL technical report. arXiv preprint arXiv:2502.13923 (2025)

  5. [4]

    In: CVPR

    Barron, J.T., Mildenhall, B., Verbin, D., Srinivasan, P.P., Hedman, P.: Mip-NeRF 360: Unbounded anti-aliased neural radiance fields. In: CVPR. pp. 5470–5479 (2022)

  6. [5]

    In: CVPR

    Brahmbhatt, S., Hays, J.: DeepNav: Learning to navigate large cities. In: CVPR. pp. 5193–5202 (2017)

  7. [6]

    In: CVPR

    Chen, H., Suhr, A., Misra, D., Snavely, N., Artzi, Y.: TOUCHDOWN: Natural language navigation and spatial reasoning in visual street environments. In: CVPR. pp. 12538–12547 (2019)

  8. [7]

    arXiv preprint arXiv:2507.06261 (2025)

    Comanici, G., et al.: Gemini 2.5: Pushing the frontier with advanced reason- ing, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261 (2025)

Show all 63 references
  1. [8]

    In: CVPR

    Das, A., Datta, S., Gkioxari, G., Lee, S., Parikh, D., Batra, D.: Embodied question answering. In: CVPR. pp. 1–10 (2018)

  2. [9]

    In: CoRL

    Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., Koltun, V.: CARLA: An open urban driving simulator. In: CoRL. vol. 78, pp. 1–16 (2017)

  3. [10]

    TETCI6(2), 230–244 (2022)

    Duan, J., Yu, S., Tan, H.L., Zhu, H., Tan, C.: A survey of embodied AI: From simulators to research tasks. TETCI6(2), 230–244 (2022)

  4. [11]

    Feng, J., Zhang, J., Liu, T., Zhang, X., Ouyang, T., Yan, J., Du, Y., Guo, S., Li, Y.: CityBench: Evaluating the capabilities of large language models for urban tasks. In: KDD. pp. 5413–5424 (2025)

  5. [12]

    arXiv preprint arXiv:2410.09604 (2024)

    Gao, C., Zhao, B., Zhang, W., Mao, J., Zhang, J., Zheng, Z., Man, F., Fang, J., Zhou, Z., Cui, J., et al.: EmbodiedCity: A benchmark platform for embodied agent in real-world city environment. arXiv preprint arXiv:2410.09604 (2024)

  6. [13]

    In: CVPR

    Haas,L.,Skreta,M.,Alberti,S.,Finn,C.:PIGEON:Predictingimagegeolocations. In: CVPR. pp. 12893–12902 (2024)

  7. [14]

    In: NeurIPS

    Hong, Y., Sun, R., Li, B., Yao, X., Wu, M., Chien, A., Yin, D., Wu, Y.N., Wang, Z., Chang, K.W.: Embodied web agents: Bridging physical-digital realms for inte- grated agent intelligence. In: NeurIPS. vol. 38 (2025)

  8. [15]

    9118–9147 (2022)

    Huang, W., Abbeel, P., Pathak, D., Mordatch, I.: Language models as zero-shot planners:Extractingactionableknowledgeforembodiedagents.In:ICML.vol.162, pp. 9118–9147 (2022)

  9. [16]

    In: CoRL

    Ichter, B., Brohan, A., Chebotar, Y., Finn, C., Hausman, K., Herzog, A., Ho, D., Ibarz, J., Irpan, A., Jang, E., Julian, R., Kalashnikov, D., Levine, S., Lu, Y., Parada, C., Rao, K., Sermanet, P., Toshev, A.T., Vanhoucke, V., Xia, F., Xiao, T., Xu, P., Yan, M., Brown, N., Ahn,...

  10. [17]

    AAAI40(22), 18342–18350 (2026)

    Ji, Y., Zhu, Z., Zhao, Y., Liu, B., Gao, C., Zhao, Y., Qiu, S., Hu, Y., Yin, Q.: To- wards autonomous UAV visual object search in city space: Benchmark and agentic methodology. AAAI40(22), 18342–18350 (2026)

  11. [18]

    In: ICRA

    Jiao, J., He, J., Liu, C., Aegidius, S., Hu, X., Braud, T., Kanoulas, D.: LiteVLoc: Map-lite visual localization for image goal navigation. In: ICRA. pp. 5244–5251 (2025)

  12. [19]

    IEEE Access13, 162467–162504 (2025)

    Kawaharazuka, K., Oh, J., Yamada, J., Posner, I., Zhu, Y.: Vision-language-action models for robotics: A review towards real-world applications. IEEE Access13, 162467–162504 (2025)

  13. [20]

    ACM TOG42(4), 139 (2023)

    Kerbl, B., Kopanas, G., Leimkühler, T., Drettakis, G.: 3D gaussian splatting for real-time radiance field rendering. ACM TOG42(4), 139 (2023)

  14. [21]

    In: CVPR

    Khanna, M., Ramrakhya, R., Chhablani, G., Yenamandra, S., Gervet, T., Chang, M., Kira, Z., Chaplot, D.S., Batra, D., Mottaghi, R.: GOAT-Bench: A benchmark for multi-modal lifelong navigation. In: CVPR. pp. 16373–16383 (2024)

  15. [22]

    In: ICCV

    Lee, J., Miyanishi, T., Kurita, S., Sakamoto, K., Azuma, D., Matsuo, Y., Inoue, N.: CityNav: A large-scale dataset for real-world aerial navigation. In: ICCV. pp. 5912–5922 (2025)

  16. [23]

    AAAI38(17), 18517–18526 (2024)

    Li, J., Padmakumar, A., Sukhatme, G., Bansal, M.: VLN-Video: Utilizing driv- ing videos for outdoor vision-and-language navigation. AAAI38(17), 18517–18526 (2024)

  17. [24]

    In: NeurIPS

    Li, M., Zhao, S., Wang, Q., Wang, K., Zhou, Y., Srivastava, S., Gokmen, C., Lee, T., Li, L.E., Zhang, R., Liu, W., Liang, P., Fei-Fei, L., Mao, J., Wu, J.: Embodied Agent Interface: Benchmarking LLMs for embodied decision making. In: NeurIPS. vol. 37, pp. 100428–100534 (2024)

  18. [25]

    In: CVPR

    Liu, X., Li, J., Jiang, Y., Sujay, N., Yang, Z., Zhang, J., Abanes, J., Zhang, J., Feng, C.: CityWalker: Learning embodied urban navigation from web-scale videos. In: CVPR. pp. 6875–6885 (2025)

  19. [26]

    IEEE/ASME TMECH30(6), 7253–7274 (2025)

    Liu, Y., Chen, W., Bai, Y., Liang, X., Li, G., Gao, W., Lin, L.: Aligning cyber space with physical world: A comprehensive survey on embodied AI. IEEE/ASME TMECH30(6), 7253–7274 (2025)

  20. [27]

    arXiv preprint arXiv:2606.16953 (2026)

    Liu,Z.,He,H.,Alumootil,V.,Pandya,A.,Squicciarini,B.,Wu,W.,Zhou,B.:Side- walkBench: Benchmarking visual navigation on urban sidewalks. arXiv preprint arXiv:2606.16953 (2026)

  21. [28]

    In: NeurIPS

    Majumdar, A., Aggarwal, G., Devnani, B., Hoffman, J., Batra, D.: ZSON: Zero- shot object-goal navigation using multimodal goal embeddings. In: NeurIPS. vol. 35, pp. 32340–32352 (2022)

  22. [29]

    In: ECCV

    Mildenhall, B., Srinivasan, P.P., Tancik, M., Barron, J.T., Ramamoorthi, R., Ng, R.: NeRF: Representing scenes as neural radiance fields for view synthesis. In: ECCV. vol. 12346, pp. 405–421 (2020)

  23. [30]

    In: NeurIPS

    Mirowski, P., Grimes, M., Malinowski, M., Hermann, K.M., Anderson, K., Teplyashin, D., Simonyan, K., Kavukcuoglu, K., Zisserman, A., Hadsell, R.: Learn- ing to navigate in cities without a map. In: NeurIPS. vol. 31 (2018)

  24. [31]

    arXiv preprint arXiv:2506.01952 (2025)

    Miyai, A., Zhao, Z., Egashira, K., Sato, A., Sunada, T., Onohara, S., Yamanishi, H., Toyooka, M., Nishina, K., Maeda, R., et al.: WebChoreArena: Evaluating web browsing agents on realistic tedious web tasks. arXiv preprint arXiv:2506.01952 (2025)

  25. [32]

    In: NAACL

    Mukhopadhyay, S., Rajgaria, A., Khatiwada, P., Shrivastava, M., Roth, D., Gupta, V.: MAPWise: Evaluating vision-language models for advanced map queries. In: NAACL. pp. 9348–9378 (2025)

  26. [33]

    OpenAI: GPT-5 system card. Tech. rep., OpenAI (2025) 18 K. Watanabe et al

  27. [34]

    In: EMNLP-IJCNLP

    Paz-Argaman, T., Tsarfaty, R.: RUN through the streets: A new dataset and base- line models for realistic urban navigation. In: EMNLP-IJCNLP. pp. 6449–6455 (2019)

  28. [35]

    In: ICLR (2024)

    Puig, X., Undersander, E., Szot, A., Cote, M.D., Yang, T.Y., Partsey, R., Desai, R., Clegg, A., Hlavac, M., Min, S.Y., Vondruš, V., Gervet, T., Berges, V.P., Turner, J.M., Maksymets, O., Kira, Z., Kalakrishnan, M., Malik, J., Chaplot, D.S., Jain, U., Batra, D., Rai, A., Mottag...

  29. [36]

    arXiv preprint arXiv:1712.03931 (2017)

    Savva,M.,Chang,A.X.,Dosovitskiy,A.,Funkhouser,T.,Koltun,V.:MINOS:Mul- timodal indoor simulator for navigation in complex environments. arXiv preprint arXiv:1712.03931 (2017)

  30. [37]

    In: ICCV

    Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., Parikh, D., Batra, D.: Habitat: A platform for embodied AI research. In: ICCV. pp. 9339–9347 (2019)

  31. [38]

    In: ACL-IJCNLP

    Schumann, R., Riezler, S.: Generating landmark navigation instructions from maps as a graph-to-text problem. In: ACL-IJCNLP. pp. 489–502 (2021)

  32. [39]

    In: CoRL

    Shah, D., Sridhar, A., Dashora, N., Stachowicz, K., Black, K., Hirose, N., Levine, S.: ViNT: A foundation model for visual navigation. In: CoRL. vol. 229, pp. 711– 733 (2023)

  33. [40]

    In: ACM MM

    Sugimoto, N., Ebine, Y., Aizawa, K.: Building Movie Map – a tool for exploring areas in a city – and its evaluations. In: ACM MM. pp. 3330–3338 (2020)

  34. [42]

    Wang, S., Liang, C., Gao, Y., Yu, E., Li, S., Li, J., Wang, H.: CitySeeker: How do VLMs explore embodied urban navigation with implicit human needs? In: ICLR (2026)

  35. [43]

    arXiv preprint arXiv:2508.18265 (2025)

    Wang, W., Gao, Z., Gu, L., Pu, H., Cui, L., Wei, X., Liu, Z., Jing, L., Ye, S., Shao, J., Wang, Z., Chen, Z., Zhang, H., Yang, G., Wang, H., Wei, Q., Yin, J., Li, W., Cui, E., Chen, G., Ding, Z., Tian, C., Wu, Z., Xie, J., Li, Z., Yang, B., Duan, Y., Wang, X., Hou, Z., Hao, H....

  36. [44]

    In: ICLR (2025)

    Wu, W., He, H., He, J., Wang, Y., Duan, C., Liu, Z., Li, Q., Zhou, B.: MetaUrban: An embodied AI simulation platform for urban micromobility. In: ICLR (2025)

  37. [45]

    In: CVPR

    Xie, Z., Liu, Z., Peng, Z., Wu, W., Zhou, B.: Vid2Sim: Realistic and interactive simulation from video for urban navigation. In: CVPR. pp. 1581–1591 (2025)

  38. [46]

    Xing, S., Sun, Z., Xie, S., Chen, K., Huang, Y., Wang, Y., Li, J., Song, D., Tu, Z.: Can large vision language models read maps like a human? arXiv preprint arXiv:2503.14607 (2025)

  39. [47]

    In: ECCV

    Xu, Q., Yi, X., Xu, J., Tao, W., Ong, Y.S., Zhang, H.: Few-shot NeRF by adaptive rendering loss regularization. In: ECCV. vol. 15124, pp. 125–142 (2024)

  40. [48]

    In: CVPR

    Yang, J., Pavone, M., Wang, Y.: FreeNeRF: Improving few-shot neural rendering with free frequency regularization. In: CVPR. pp. 8254–8263 (2023)

  41. [49]

    In: ECCV

    Yang,J., Ding,R.,Brown,E., Qi,X.,Xie, S.:V-IRL:Groundingvirtualintelligence in real life. In: ECCV. vol. 15103, pp. 36–55 (2024) 360CityArena 19

  42. [50]

    In: CVPR

    Yang, J., Yang, S., Gupta, A.W., Han, R., Fei-Fei, L., Xie, S.: Thinking in space: How multimodal large language models see, remember, and recall spaces. In: CVPR. pp. 10632–10643 (2025)

  43. [51]

    In: ICML

    Yang, R., Chen, H., Zhang, J., Zhao, M., Qian, C., Wang, K., Wang, Q., Ko- ripella, T.V., Movahedi, M., Li, M., Ji, H., Zhang, H., Zhang, T.: EmbodiedBench: Comprehensive benchmarking multi-modal large language models for vision-driven embodied agents. In: ICML. vol. 267, pp. ...

  44. [52]

    IJCV129, 2136–2174 (2021)

    Zaffar, M., Garg, S., Milford, M., Kooij, J., Flynn, D., McDonald-Maier, K., Ehsan, S.:VPR-Bench:Anopen-sourcevisualplacerecognitionevaluationframeworkwith quantifiable viewpoint and appearance change. IJCV129, 2136–2174 (2021)

  45. [53]

    In: ECCV

    Zhang, D., Wang, C., Wang, W., Li, P., Qin, M., Wang, H.: Gaussian in the wild: 3D Gaussian splatting for unconstrained image collections. In: ECCV. vol. 15134, pp. 341–359 (2024)

  46. [54]

    In: CoRL

    Zhang, M., Qu, K., Patil, V., Cadena, C., Hutter, M.: Tag Map: A text-based map for spatial reasoning and navigation with large language models. In: CoRL. vol. 270, pp. 2120–2146 (2024)

  47. [55]

    In: ICLR (2024)

    Zhou, S., Xu, F.F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., et al.: WebArena: A realistic web environment for building autonomous agents. In: ICLR (2024)

  48. [56]

    In: CoRL

    Zitkovich, B., Yu, T., Xu, S., Xu, P., Xiao, T., Xia, F., Wu, J., Wohlhart, P., Welker, S., Wahid, A., Vuong, Q., Vanhoucke, V., Tran, H., Soricut, R., Singh, A., Singh, J., Sermanet, P., Sanketi, P.R., Salazar, G., Ryoo, M.S., Reymann, K., Rao, K., Pertsch, K., Mordatch, I., ...

  49. [57]

    Think: Analyze the current state and decide what to do next

  50. [58]

    Action: Choose one of the following actions: - W: move forward - LEFT/RIGHT/UP: if you see red arrows, you should select one of them. - S: turn camera to the direction of travel - Q: rotate camera upward (look around) - E: rotate camera downward (look around) - A: rotate camer...

  51. [59]

    Observation: You will receive the result of your action You will receive two types of images:

  52. [60]

    Camera view: The first-person view of what you can see in the city

  53. [61]

    thought":

    Map view (when available): A top-down map showing your current location with a red arrow indicating your position and direction 22 K. Watanabe et al. Use both images to make better navigation decisions. The map can help you understand your location and plan your route more eff...

  54. [62]

    Turn left at the first intersection

  55. [63]

    Object” isreplacedwithatask-specificobjectnamesuchas“vend- ing machine

    Stop in front of Surugaya Purchase Center. Vision Language Navigation Task Prompt Your task is to follow the directions to reach your destination. Please follow the instructions below: {Directions} 28 K. Watanabe et al. Once you have reached your destination, output the ANSWER...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.