Pith. sign in

REVIEW 4 major objections 5 minor 2 cited by

Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective

T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read A neuroscience-derived six-module pipeline could give AI agents human-like spatial reasoning.

desk verdict A useful organizing survey with a reasonable six-module taxonomy, but the 'first work' claim oversells the novelty and the framework is asserted rather than derived. read the letter →

arxiv 2509.09154 v1 pith:CAG5ZFYM submitted 2025-09-11 cs.AI cs.CV

classification cs.AIcs.CV
keywords spatialreasoningagenticAIneuroscience-inspiredcognitivemapmemorymultisensoryintegrationegocentric-allocentricconversionintelligence
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the reason agentic AI systems still reason poorly about space is architectural: they process symbols and language, not the integrated multisensory, map-like representations humans use. Drawing on neuroscience findings about how the brain perceives, integrates, remembers, and reasons about space, it proposes a six-module computational framework—bio-inspired multimodal sensing, multi-sensory integration, egocentric–allocentric conversion, an artificial cognitive map, spatial memory, and spatial reasoning—as a blueprint for building spatially intelligent agents. The paper then uses this framework as an evaluation lens to review recent methods, benchmarks, and applications, identifying specific gaps such as missing landmark anchoring, drift correction, context remapping, and bidirectional perspective shifts. A sympathetic reader would take away that this is a proposal and roadmap: a structured way to organize research toward human-like spatial intelligence, not yet an implemented system.

What carries the argument

The central object is the six-module computational framework itself, presented as a perspective landscape for agentic spatial intelligence. Its load-bearing components are the artificial cognitive map, built from a grid-cells layer (hexagonal metric encoding, path integration, landmark anchoring) and a place-cells layer (topological graph, contextual remapping, memory indexing, prospective coding), and the spatial neural memory module (semantic–spatial encoding, episodic memory with compression, adaptive updating). The framework also organizes spatial reasoning behaviors into three tiers—perceptual inference, hidden-state inference, and policy selection—which then serves as the taxonomy for

What would settle it

An ablation experiment: build an agent implementing all six modules, then remove each module one at a time and test on novel-view perspective taking and long-horizon navigation. If removing the cognitive map or spatial memory does not degrade performance, the claim that these modules are essential components of spatial reasoning is falsified. A reader could also check neuroscience: if spatial behavior is shown to rely on a single non-hierarchical mechanism, the framework's biological premise fails.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central claim is that human-like spatial intelligence in AI agents can be engineered by transposing the functional organization of human spatial cognition into six computation modules, and that this is the first neuroscience-grounded framework design for agentic spatial reasoning. The modules form a perception–cognition–action pipeline: multimodal sensory input; an information processing module that calibrates, denoises, attention-gates, and fuses signals; an egocentric-to-allocentric conversion that builds viewpoint-independent 3D maps; a cognitive map with grid-cell-like metric layers and place-cell-like topological layers; a spatial neural memory with semanti

Load-bearing premise

The load-bearing premise is that human spatial cognition can be decomposed into six sequential modules and that instantiating those modules in an AI system will produce human-like spatial reasoning; the paper offers no evidence that this set of modules is necessary, sufficient, or correctly ordered.

Editorial extensions

If this is right

  • Agents built from the six modules should generalize spatial reasoning to new and unstructured environments better than current vision-language pipelines, because the framework supplies the missing internal 3D map and memory.
  • The framework-guided analysis implies that today's key bottlenecks are not model scale but missing landmark anchoring, drift correction, contextual remapping, and bidirectional perspective conversion.
  • Benchmarks should be reorganized by cognitive level—perceptual inference, hidden-state inference, policy selection—to expose which spatial abilities agents actually lack.
  • Future systems should combine predictive world models with explicit multistep spatial reasoning, rather than treating spatial questions as pure language tasks.
  • Deployment of spatial reasoning on edge and neuromorphic hardware is a necessary part of the roadmap if agents are to act in real time.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the framework is right, current spatial failures of large vision-language models, such as perspective-taking hallucination, are primarily architecture failures rather than scale failures; adding explicit geometric memory should matter more than adding parameters.
  • A testable extension: instantiate the six modules with existing SLAM, 3D-VLM, graph memory, and model-based reinforcement learning, then run a six-module versus end-to-end comparison on viewpoint-transfer and long-horizon tasks; an ablation of the cognitive map would test its necessity.
  • The framework suggests a new benchmark design principle: tasks should be labeled by the cognitive module they stress—sensing, integration, conversion, map, memory, reasoning—enabling modular diagnosis of agent failures.
  • The neuroscience mapping is a working hypothesis; if human spatial reasoning turns out not to be modular in this sequence, the framework would still survive as an engineering heuristic but would lose its biological grounding.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a neuroscience-inspired framework for agentic spatial intelligence, decomposing human spatial cognition into six computation modules: bio-inspired multimodal sensing, multi-sensory integration, egocentric–allocentric conversion, an artificial cognitive map, spatial neural memory, and spatial reasoning. These modules are presented as a serial pipeline in Algorithm 5, grounded in a review of spatial cognition neuroscience (Sec. 2.1), and then used as the lens for a survey of recent methods (Sec. 3.1), a categorization of benchmarks (Sec. 3.2.1), and a set of future research directions (Sec. 4). The paper claims this is the first neuroscience-based framework design for agentic spatial reasoning.

Significance. If established, the framework could give the field a useful common vocabulary and a structured research agenda, and the survey elements are timely: the neuroscience summaries in Sec. 2.1 are broad and readable, Tables 3–7 organize a large body of recent work, and the benchmark categorization in Table 8 is a useful resource. The pseudo-code algorithms, though illustrative, make the proposal concrete. However, the central claim is not validated: the six-module decomposition is asserted rather than derived, and the paper's novelty and gap analyses depend on that unsupported decomposition. The strength of the paper is therefore as a perspective and organizing taxonomy, not as an established foundation.

major comments (4)
  1. [Sec. 2.2, Algorithm 5 (lines 8–11), and Contributions] The six modules are called 'essential' and are arranged as a strictly sequential pipeline, but no argument establishes necessity, sufficiency, or order. This is not a local presentation issue: Sec. 3.1 derives RG1–RG5 by evaluating methods against this module set, and the conclusion states the framework 'pav[es] the way to achieve the human spatial intelligence in agentic systems.' The manuscript itself supplies a counterexample to the serial flow: Sec. 2.2.2 says scene abstraction 'maintains bidirectional connectivity with the cognitive map module,' yet Algorithm 5 calls EgocentricAllocentricConversion and InternalMemoryModel serially and never routes information back to the egocentric module. The authors should either justify the decomposition (e.g., from a stated design criterion or comparative cognitive evidence) or explicitly reframe it as one possible organizing taxonomy, and recon
  2. [Sec. 2.2.5, Eq. (6)] The paper states that Hierarchical Active Inference is 'widely regarded as a minimal model for studying spatial reasoning' and uses it to categorize spatial reasoning into 3D Perceptual Inference, Hidden-State Inference, and Policy Selection. No citation or derivation supports 'minimal,' and the mapping from Eq. (6) to the three categories is not shown; it is an interpretive choice. The same taxonomy later organizes the benchmark analysis in Sec. 3.2.1, so an unsupported categorization propagates into the evaluation structure. Please either derive the categories from HAI and cite the relevant literature, or weaken the claim and describe the taxonomy as a proposed schema whose validity remains to be tested.
  3. [Sec. 3.1, RG1–RG5] The framework-guided gap analysis is in part self-supporting, which is a correctness risk: RG3, for example, identifies that current grid/place-cell models 'lack landmark anchoring, drift correction, multi-field coding, and context-dependent remapping'—precisely the components of the proposed Cognitive Map Module. The analysis can therefore organize the literature without providing evidence that the framework's modules are the right ones. A concrete test would be to show that methods strong on framework-defined components outperform others on benchmarks that are not themselves derived from the framework, or to show that the proposed decomposition predicts specific failure modes in existing agents. Absent such a test, the claim that this is 'the first work to explore the neuroscience-based framework design for agentic spatial reasoning' is broader than the evidence supports.
  4. [Sec. 1, Related Works and Contributions] The novelty claim—'to the best of our knowledge, this is the first work to explore the neuroscience-based framework design for agentic spatial reasoning'—is presented without a systematic comparison with prior neuroscience-inspired perspectives that the paper itself lists, such as Refs. [125], [129], and [166]. Those works also draw on neuroscience to discuss embodied agents and reasoning. The authors should either delineate precisely what the six-module decomposition and framework-guided gap analysis add over these antecedents, or soften the 'first' claim. The comparison matters because this headline contribution is used throughout the paper as a mark of significance.
minor comments (5)
  1. [Sec. 4.3 and Sec. 4.5] Cross-reference errors: Sec. 4.5 cites 'RG-5 in Sec. 3.1.4' but RG5 is defined in Sec. 3.1.5; Sec. 4.3 cites 'RG-3 in Sec. 2.2' but the gap analysis for the cognitive map appears in Sec. 3.1.3. Please fix.
  2. [Throughout] Several typos and inconsistencies should be corrected: '3D Contrusction' (Sec. 3.1.2), 'scence abstraction and persspective change' (Fig. 9), 'pattern seperation' (Sec. 2.1.1), 'Key Feautres' (Table 2), 'ARKitScences' (Table 8), 'A VFormer' (Sec. 3.1.1), 'Parallely' (Sec. 4.1), and 'generaive' (Sec. 4.5).
  3. [Algorithms 1–5] The pseudo-code uses functions such as InformationProcessing, EgocentricAllocentricConversion, InternalMemoryModel, and Reasoning that are not formally specified. This is acceptable for a conceptual perspective, but the paper should explicitly label the algorithms as illustrative or define the high-level interfaces, to avoid giving the impression of a runnable specification.
  4. [Algorithm 3] The input list says 'Require: unified_latent_space, allocentric_map,' but line 10 uses 'trajectory,' which is neither an input nor defined before use. Please clarify the source of the trajectory variable (e.g., from working memory, as in Algorithm 2).
  5. [Table 3] The venue for RegBN [61] is listed as 'ANIPS' but should be 'NeurIPS.' Also, in Table 8, ARKitScenes is misspelled as 'ARKitScences'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the framework is a proposal, not a derivation, and self-citations are non-load-bearing.

full rationale

This is a perspective/survey paper, not a derivation with fitted parameters or predictions. The six-module framework is presented as a conceptual proposal grounded in prior neuroscience literature, and the subsequent framework-guided gap analysis is explicitly normative: gaps are defined as absences of modules the authors argue are essential. That is a transparent evaluation lens, not a case where an output is equivalent to an input by construction. No equation in the paper is fitted to data and then renamed as a prediction; the equations reproduced (Bayes, free energy, successor representation) are standard textbook definitions cited to external sources. Self-citations appear (e.g., [18], [19], [36], possibly [129]), but they are used as background, examples of techniques, or modular components; none is load-bearing for the paper's central claim that a neuroscience-inspired framework for agentic spatial reasoning is useful. The internal tension between the bidirectional connectivity described in Sec. 2.2.2 and the strictly serial Algorithm 5 is a consistency/correctness concern, not circularity. The novelty claim ('first work') is an assertion unsupported by a systematic literature search, but that is a scope/evidence concern, not a circularity. Overall, there is no specific reduction of a conclusion to its own premises by definition or by self-citation chain, so the paper should not be scored as circular.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No free parameters and no invented physical entities; the paper is conceptual. The central claim rests on unvalidated domain assumptions, especially that spatial cognition decomposes into the six proposed modules and that Hierarchical Active Inference is the right organizing taxonomy.

assumptions (5)
  • domain assumption The brain processes spatial information through a hierarchical flow from multisensory perception to cognitive map to reasoning (Section 2.1.1).
    This hierarchy is presented as the backbone for the proposed framework; the six modules mirror this flow.
  • domain assumption Grid cells in MEC provide Euclidean metric and place cells in HPC provide topological context, together forming the cognitive map (Section 2.1.2).
    The cognitive map module's grid cells layer and place cells layer depend on this theory, which is widely cited but still an active area of debate.
  • ad hoc to paper Hierarchical Active Inference is a minimal model for spatial reasoning and justifies the three behavior categories (perceptual inference, hidden-state inference, policy selection) in Section 2.2.5.
    The paper adopts HAI as the organizing lens for spatial reasoning behaviors without arguing why it is minimal or complete.
  • ad hoc to paper Neuroscience principles can be transferred to AI modules at the chosen level of abstraction (Section 2.2).
    The framework assumes functional equivalence between brain regions and algorithmic modules; no formal or empirical justification is provided.
  • domain assumption The surveyed methods in Section 3.1 are representative of the field.
    The selection is narrative and framework-guided; no systematic search or inclusion criteria are reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective." pith.science (2026). https://pith.science/paper/CAG5ZFYM

@misc{pith2026250909154,
  author       = {Pith},
  title        = {Pith review of: Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CAG5ZFYM}},
  note         = {Machine review of arXiv:2509.09154}
}
read the original abstract

Recent advances in agentic AI have led to systems capable of autonomous task execution and language-based reasoning, yet their spatial reasoning abilities remain limited and underexplored, largely constrained to symbolic and sequential processing. In contrast, human spatial intelligence, rooted in integrated multisensory perception, spatial memory, and cognitive maps, enables flexible, context-aware decision-making in unstructured environments. Therefore, bridging this gap is critical for advancing Agentic Spatial Intelligence toward better interaction with the physical 3D world. To this end, we first start from scrutinizing the spatial neural models as studied in computational neuroscience, and accordingly introduce a novel computational framework grounded in neuroscience principles. This framework maps core biological functions to six essential computation modules: bio-inspired multimodal sensing, multi-sensory integration, egocentric-allocentric conversion, an artificial cognitive map, spatial memory, and spatial reasoning. Together, these modules form a perspective landscape for agentic spatial reasoning capability across both virtual and physical environments. On top, we conduct a framework-guided analysis of recent methods, evaluating their relevance to each module and identifying critical gaps that hinder the development of more neuroscience-grounded spatial reasoning modules. We further examine emerging benchmarks and datasets and explore potential application domains ranging from virtual to embodied systems, such as robotics. Finally, we outline potential research directions, emphasizing the promising roadmap that can generalize spatial reasoning across dynamic or unstructured environments. We hope this work will benefit the research community with a neuroscience-grounded perspective and a structured pathway. Our project page can be found at Github.

Figures

Figures reproduced from arXiv: 2509.09154 by the authors.

Figure 1
Figure 1. Illustration of neuroscience-inspired Agentic Spatial Intelligence. As the core functions of human [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. The neuroscience-based cognitive map. It is rooted in the hippocampus (orange) and entorhinal cortex [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Memory systems in human cognition: working, episodic, and long-term. Episodic Memory. Tulving [193] defined episodic memory as the system that encodes, stores, and re￾trieves autobiographical experiences within specific spatiotemporal contexts. It supports episodic recollec￾tion by enabling the mental recall of the experienced events with both spatial and temporal contexts (e.g., “where and when they occurred”). Bio… view at source ↗
Figures from the paper (12 more)
Figure 4
Figure 4. Figure 4: Development of the backbone neuroscience models for spatial reasoning. [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: The architecture of TEM. (A) Generative model showing the top-down process from actions ( [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: The proposed framework of Agentic Spatial Intelligence. Following human cognition from perception, [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: The multi-sensory input required for agents. It consists of bio-inspired modalities: visual, auditory, [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]
Figure 8
Figure 8. Figure 8: Information Processing Module (IPM) for Spatial Reasoning. Multisensory inputs are preprocessed and [PITH_FULL_IMAGE:figures/full_fig_p015_8.png]
Figure 9
Figure 9. Figure 9: Egocentric-to-Allocentric transformation. It commence by projecting previsous cross-modal latent [PITH_FULL_IMAGE:figures/full_fig_p017_9.png]
Figure 10
Figure 10. Figure 10: Cognitive map for spatial reasoning. This module encodes the agent’s internal representation via a [PITH_FULL_IMAGE:figures/full_fig_p019_10.png]
Figure 11
Figure 11. Figure 11: (A) Spatial–Semantic Encoding. Sensory inputs are transformed into spatial and semantic repre￾sentations for unified understanding across modalities. (B) Episodic Spatial Memory. Multimodal inputs are encoded by a transformer, compressed via entropy modeling for effic…
Figure 12
Figure 12. Figure 12: Reasoning Module for Spatial Reasoning. Predictive world modelling refers to the predictive ability [PITH_FULL_IMAGE:figures/full_fig_p022_12.png]
Figure 13
Figure 13. Figure 13: Architecture from [45]. Although the network achieves selective visual representation for improved generalization, the codebook bottleneck limits adaptability to novel environments and restrict the expressive￾ness of learned features, especially in dynamic or highly v…
Figure 14
Figure 14. Figure 14: The representative works tailored in (a) cognitive map module and (b) spatial neural memory module [PITH_FULL_IMAGE:figures/full_fig_p033_14.png]
Figure 15
Figure 15. Figure 15: Applications from Agentic Spatial Intelligence, including (a) Virtual and (b) Physical Applications. [PITH_FULL_IMAGE:figures/full_fig_p041_15.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EAGOR: Embodied Reasoning in Omni-direction

    cs.RO 2026-07 conditional novelty 7.0 of 10

    EAGOR reformulates embodied 360-degree directional reasoning as recursive Bayesian estimation on a spherical manifold using spherical harmonics, achieving training-free, rotation-equivariant target tracking.

  2. VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation

    cs.RO 2026-07 conditional novelty 5.5 of 10

    VistaVLA lifts multi-view vision-language features into 3D Gaussians, compresses them 99% via Merge-then-Query, and improves real-robot manipulation success by ~23% over baselines.

Reference graph

Works this paper leans on

231 extracted references · 3 canonical work pages · cited by 2 Pith papers

  1. [125]

    Jian Liu, Xiongtao Shi, Thai Duy Nguyen, Haitian Zhang, et al . 2025. Neural Brain: A Neuroscience-inspired Framework for Embodied Agents.arXiv preprint arXiv:2505.07634(2025)

  2. [129]

    Zinan Liu, Haoran Li, Jingyi Lu, et al. 2025. Nature’s Insight: A Novel Framework and Comprehensive Analysis of Agentic Reasoning Through the Lens of Neuroscience.arXiv preprint arXiv:2505.05515(2025)

  3. [166]

    Rizwan Qureshi, Ranjan Sapkota, Abbas Shah, Amgad Muneer, Anas Zafar, Ashmal Vayani, et al. 2025. Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact.arXiv preprint arXiv:2507.00951(2025)

  4. [1]

    Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al. 2022. Do as i can, not as i say: Grounding language in robotic affordances.arXiv preprint arXiv:2204.01691(2022)

  5. [2]

    Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong. 2021. Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text.Advances in neural information processing systems34 (2021), 24206–24221

  6. [3]

    Fayçal Aït Aoudia, Jakob Hoydis, Merlin Nimier-David, Sebastian Cammerer, and Alexander Keller. 2025. Sionna RT: Technical Report.arXiv preprint arXiv:2504.21719(2025)

  7. [4]

    Iro Armeni, Zhi-Yang He, JunYoung Gwak, Amir R Zamir, Martin Fischer, Jitendra Malik, and Silvio Savarese. 2019. 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and Camera. InProceedings of the IEEE International Conference on Computer Vision. 5664–5673. Preprint (September 2025). 46 Manh, et al

  8. [5]

    Armen Avetisyan, Manuel Dahnert, Angela Dai, Manolis Savva, Angel X Chang, and Matthias Nießner. 2019. Scan2cad: Learning cad model alignment in rgb-d scans. InProceedings of the IEEE/CVF Conference on computer vision and pattern recognition. 2614–2623

Show all 231 references
  1. [6]

    Andrea Banino, Caswell Barry, Benigno Uria, Charles Blundell, Timothy Lillicrap, Piotr Mirowski, Alexander Pritzel, Martin J Chadwick, Thomas Degris, Joseph Modayil, et al. 2018. Vector-based navigation using grid-like representa- tions in artificial agents.Nature557, 7705 (20...

  2. [7]

    Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, Yuri Feigin, Peter Fu, Thomas Gebauer, et al. 2021. ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data. InThirty-fifth Conference on Neural Information Processing Systems Datasets and...

  3. [8]

    Peter W Battaglia, Jessica B Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, et al. 2018. Relational inductive biases, deep learning, and graph networks.arXiv preprint arXiv:180...

  4. [9]

    Tristan Baumann and Hanspeter A Mallot. 2023. Metric information in cognitive maps: Euclidean embedding of non-Euclidean environments.PLOS Computational Biology19, 12 (2023), e1011748

  5. [10]

    Michael S Beauchamp, Kathryn E Lee, Brenna D Argall, and Alex Martin. 2004. Integration of auditory and visual information about objects in superior temporal sulcus.Neuron41, 5 (2004), 809–823

  6. [11]

    Raunaq Bhirangi, Tess Hellebrekers, Carmel Majidi, and Abhinav Gupta. 2022. ReSkin: versatile, replaceable, lasting tactile skins. InConference on Robot Learning. 587–597

  7. [12]

    Mahtab Bigverdi, Zelun Luo, Cheng-Yu Hsieh, Ethan Shen, Dongping Chen, Linda G Shapiro, and Ranjay Krishna

  8. [13]

    Jennifer K Bizley and Yale E Cohen. 2013. The what, where and how of auditory-object perception.Nature Reviews Neuroscience14, 10 (2013), 693–707. doi:10.1038/nrn3565

  9. [14]

    Florian Bordes, Quentin Garrido, Justine T Kao, Adina Williams, Michael Rabbat, and Emmanuel Dupoux. 2025. IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments.arXiv preprint arXiv:2506.09849(2025)

  10. [15]

    Randy L Buckner, Jessica R Andrews-Hanna, and Daniel L Schacter. 2008. The brain’s default network: anatomy, function, and relevance to disease.Annals of the new York Academy of Sciences1124, 1 (2008), 1–38

  11. [16]

    Neil Burgess. 2008. Spatial cognition and the brain.Annals of the New York Academy of Sciences1124, 1 (2008), 77–97

  12. [17]

    Patrick Byrne, Suzanna Becker, and Neil Burgess. 2007. Remembering the past and imagining the future: a neural model of spatial memory and imagery.Psychological review114, 2 (2007), 340

  13. [18]

    Erik Cambria, Rui Mao, Melvin Chen, Zhaoxia Wang, and Seng-Beng Ho. 2023. Seven Pillars for the Future of Artificial Intelligence.IEEE Intelligent Systems38, 6 (2023), 62–69

  14. [19]

    Erik Cambria, Rui Mao, Xulang Zhang, Luwei Xiao, Tiesunlong Shen, and Avinash Anand. 2026. SenticNet 9: Generative Commonsense for Emotion AI via Conceptual Primitive Discovery and Time Shift Mechanism.IEEE Transactions on Computational Social Systems13 (2026)

  15. [20]

    Sebastian Cammerer, Guillermo Marcus, Tobias Zirr, Fayçal Aït Aoudia, et al. 2025. Sionna Research Kit: A GPU- Accelerated Research Platform for AI-RAN.arXiv preprint arXiv:2505.15848(2025)

  16. [21]

    Vincent Cartillier, Zhile Ren, Neha Jain, Stefan Lee, Irfan Essa, and Dhruv Batra. 2021. Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric Views.Proceedings of the AAAI Conference on Artificial Intelligence35, 2 (2021), 964–972. doi:10.160...

  17. [22]

    Geeling Chau, Christopher Wang, Sabera Talukder, Vighnesh Subramaniam, Saraswati Soedarmadji, Yisong Yue, Boris Katz, and Andrei Barbu. 2025. Population transformer: Learning population-level representations of neural activity.ArXiv(2025), arXiv–2406

  18. [23]

    Changan Chen, Unnat Jain, Carl Schissler, Sebastia Vicenc Amengual Gari, Ziad Al-Halah, Vamsi Krishna Ithapu, Philip Robinson, and Kristen Grauman. 2020. Soundspaces: Audio-visual navigation in 3d environments. InEuropean conference on computer vision. Springer, 17–36

  19. [24]

    Chang Chen, Jiaming Zhang, Kailun Yang, Kunyu Peng, and Rainer Stiefelhagen. 2023. Trans4Map: Revisiting holistic bird’s-eye-view mapping from egocentric images to allocentric semantics with vision transformers. InProceedings of the IEEE/CVF Winter Conference on Applications o...

  20. [25]

    Shiqi Chen, Tongyao Zhu, Ruochen Zhou, Jinghan Zhang, Siyang Gao, Juan Carlos Niebles, Mor Geva, Junxian He, Jiajun Wu, and Manling Li. 2025. Why is spatial reasoning hard for vlms? an attention mechanism perspective on focus areas.arXiv preprint arXiv:2503.01773(2025)

  21. [26]

    Yinpeng Chen, DeLesley Hutchins, Aren Jansen, Andrey Zhmoginov, David Racz, and Jesper Andersen. 2024. Melodi: Exploring memory compression for long contexts.arXiv preprint arXiv:2410.03156(2024)

  22. [27]

    Hongrong Cheng, Miao Zhang, and Javen Qinfeng Shi. 2024. A survey on deep neural network pruning: Taxonomy, comparison, analysis, and recommendations.IEEE Transactions on Pattern Analysis and Machine Intelligence(2024). Preprint (September 2025). Mind Meets Space: Rethinking A...

  23. [28]

    Junmo Cho, Jaesik Yoon, and Sungjin Ahn. 2024. Spatially-Aware Transformer for Embodied Agents.arXiv preprint arXiv:2402.15160(2024)

  24. [29]

    Sanjoy Chowdhury, Sayan Nag, Subhrajyoti Dasgupta, Jun Chen, Mohamed Elhoseiny, Ruohan Gao, and Dinesh Manocha. 2024. Meerkat: Audio-visual large language model for grounding in space and time. InEuropean Conference on Computer Vision. Springer, 52–70

  25. [30]

    Nam Hoai Chu, Diep N Nguyen, Dinh Thai Hoang, et al . 2023. AI-enabled mm-Waveform configuration for autonomous vehicles with integrated communication and sensing.IEEE Internet of Things Journal10, 19 (2023), 16727–16743

  26. [31]

    Xinru Cui, Qiming Liu, Zhe Liu, and Hesheng Wang. 2024. Frontier-Enhanced Topological Memory with Improved Exploration Awareness for Embodied Visual Navigation. InEuropean Conference on Computer Vision. 296–313

  27. [32]

    Vedant Dave et al . 2024. Multimodal visual-tactile representation learning through self-supervised contrastive pre-training. In2024 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 8013–8020

  28. [33]

    Robert David, Jared Duke, Advait Jain, et al. 2021. Tensorflow lite micro: Embedded machine learning for tinyml systems.Proceedings of machine learning and systems3 (2021), 800–811

  29. [34]

    Mike Davies, Narayan Srinivasa, Tsung-Han Lin, Gautham Chinya, et al. 2018. Loihi: A neuromorphic manycore processor with on-chip learning.Ieee Micro38, 1 (2018), 82–99

  30. [35]

    Pooya Davoodi, Chul Gwon, Guangda Lai, and Trevor Morris. 2019. Tensorrt inference with tensorflow. InGPU Technology Conference

  31. [36]

    Soumyaratna Debnath, Ashish Tiwari, Kaustubh Sadekar, and Shanmuganathan Raman. 2025. RASP: Revisiting 3D Anamorphic Art for Shadow-Guided Packing of Irregular Objects. InProceedings of the Computer Vision and Pattern Recognition Conference. 5849–5858

  32. [37]

    Deepmind. 2025. Genie 3: A new frontier for world models. https://deepmind.google/discover/blog/genie-3-a-new- frontier-for-world-models/ Accessed: July. 30, 2025

  33. [38]

    Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, et al. 2023. Objaverse-xl: A universe of 10m+ 3d objects.Advances in Neural Information Processing Systems36 (2023), 35799–35813

  34. [39]

    Joseph Del Rosario, Stefano Coletta, Soon Ho Kim, Zach Mobille, Kayla Peelman, Brice Williams, et al. 2025. Lateral inhibition in V1 controls neural and perceptual contrast sensitivity.Nature Neuroscience(2025), 1–12

  35. [40]

    Mark D’Esposito and Bradley R Postle. 2015. The cognitive neuroscience of working memory.Annual review of psychology66, 1 (2015), 115–142

  36. [41]

    Christian F Doeller, Caswell Barry, and Neil Burgess. 2010. Evidence for grid cells in a human memory network. Nature463, 7281 (2010), 657–661

  37. [42]

    Yadin Dudai, Avi Karni, and Jan Born. 2015. The consolidation and transformation of memory.Neuron88, 1 (2015), 20–32

  38. [43]

    Alexander Dupuy, Jed Schwartz, Yechiam Yemini, and David Bacon. 1990. NEST: A network simulation and prototyping testbed.Commun. ACM33, 10 (1990), 63–74

  39. [44]

    Zane Durante, Qiuyuan Huang, Naoki Wake, Ran Gong, Jae Sung Park, Bidipta Sarkar, Rohan Taori, Yusuke Noda, Demetri Terzopoulos, Yejin Choi, et al. 2024. Agent ai: Surveying the horizons of multimodal interaction.arXiv preprint arXiv:2401.03568(2024)

  40. [45]

    Ainaz Eftekhar, Kuo-Hao Zeng, Jiafei Duan, Ali Farhadi, Aniruddha Kembhavi, and Ranjay Krishna. 2024. Selective Visual Representations Improve Convergence and Generalization for Embodied AI. InThe Twelfth International Conference on Learning Representations

  41. [46]

    Howard Eichenbaum. 1997. Declarative memory: Insights from cognitive neurobiology.Annual review of psychology 48, 1 (1997), 547–572

  42. [47]

    Howard Eichenbaum, Andrew P Yonelinas, and Charan Ranganath. 2007. The medial temporal lobe and recognition memory.Annu. Rev. Neurosci.30, 1 (2007), 123–152

  43. [48]

    Gary W Elko and Jürgen Meyer. 2008. Microphone arrays. InSpringer Handbook of Speech Processing, Jacob Benesty, M. Mohan Sondhi, and Yiteng Huang (Eds.). Springer, 1021–1041. doi:10.1007/978-3-540-49127-9_49

  44. [49]

    Jakob Engel, Kiran Somasundaram, Michael Goesele, Albert Sun, Alexander Gamino, Andrew Turner, et al. 2023. Project aria: A new tool for egocentric multi-modal ai research.arXiv preprint arXiv:2308.13561(2023)

  45. [50]

    Huajian Fang, Niklas Wittmer, Johannes Twiefel, Stefan Wermter, and Timo Gerkmann. 2023. Partially adaptive multichannel joint reduction of ego-noise and environmental noise. InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). ...

  46. [51]

    Kuan Fang, Alexander Toshev, Li Fei-Fei, and Silvio Savarese. 2019. Scene memory transformer for embodied agents in long-horizon tasks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 538–547

  47. [52]

    Yao Feng, Jing Lin, Sai Kumar Dwivedi, Yu Sun, Priyanka Patel, and Michael J Black. 2024. Chatpose: Chatting about 3d human pose. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2093–2103. Preprint (September 2025). 48 Manh, et al

  48. [53]

    Markus Frey, Christian F Doeller, and Caswell Barry. 2023. Probing neural representations of scene perception in a hippocampally dependent task using artificial neural networks. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2113–2121

  49. [54]

    Karl Friston. 2010. The free-energy principle: a unified brain theory?Nature Reviews Neuroscience11, 2 (2010), 127–138

  50. [55]

    Karl Friston, James Kilner, and Lee Harrison. 2006. A free energy principle for the brain.Journal of Physiology-Paris 100, 1-3 (2006), 70–87

  51. [56]

    Karl J Friston, Thomas Parr, and Bert de Vries. 2017. The graphical brain: Belief propagation and active inference. Network neuroscience1, 4 (2017), 381–414

  52. [57]

    Letian Fu, Gaurav Datta, Huang Huang, William Chung-Ho Panitch, Jaimyn Drake, Joseph Ortiz, Mustafa Mukadam, Mike Lambeta, Roberto Calandra, and Ken Goldberg. 2024. A Touch, Vision, and Language Dataset for Multimodal Alignment. InInternational Conference on Machine Learning. ...

  53. [58]

    2020.Spinnaker-a spiking neural network architecture

    Steve Furber and Petrut , Bogdan. 2020.Spinnaker-a spiking neural network architecture. Now publishers

  54. [59]

    Chuang Gan, Jeremy Schwartz, Seth Alter, Damian Mrowca, Martin Schrimpf, James Traer, et al. 2020. Threedworld: A platform for interactive multi-modal physical simulation.arXiv preprint arXiv:2007.04954(2020)

  55. [60]

    Tenenbaum

    Chuang Gan, Yiwei Zhang, Jiajun Wu, Boqing Gong, and Joshua B. Tenenbaum. 2020. Look, Listen, and Act: Towards Audio-Visual Embodied Navigation. InICRA

  56. [61]

    Morteza Ghahremani Boozandani and Christian Wachinger. 2023. Regbn: Batch normalization of multimodal data with regularization.Advances in Neural Information Processing Systems36 (2023), 21687–21701

  57. [62]

    Asif A Ghazanfar and Charles E Schroeder. 2006. Is neocortex essentially multisensory?Trends in cognitive sciences 10, 6 (2006), 278–285

  58. [63]

    Andrey Gizdov, Shimon Ullman, and Daniel Harari. 2025. Seeing More with Less: Human-like Representations in Vision Models. InProceedings of the Computer Vision and Pattern Recognition Conference. 4408–4417

  59. [64]

    P.S Goldman-Rakic. 1995. Cellular basis of working memory.Neuron14, 3 (1995), 477–485

  60. [65]

    Google. [n. d.]. Coral Edge TPU Products. Online. https://coral.ai

  61. [66]

    Google. 2025. Gemini CLI: your open-source AI agent. https://blog.google/technology/developers/introducing- gemini-cli-open-source-ai-agent/ Accessed: July. 30, 2025

  62. [67]

    Gracjan Góral, Alicja Ziarko, Michał Nauman, and Maciej Wołczyk. 2024. Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models.arXiv preprint arXiv:2409.12969(2024)

  63. [68]

    James Gornet and Matt Thomson. 2024. Automated construction of cognitive maps with visual predictive coding. Nature Machine Intelligence6, 7 (2024), 820–833

  64. [69]

    Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, et al

  65. [70]

    Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwińska, et al. 2016. Hybrid computing using a neural network with dynamic external memory.Nature538, 7626 (2016), 471–476

  66. [71]

    Madeleine Grunde-McLaughlin, Ranjay Krishna, et al. 2021. Agqa: A benchmark for compositional spatio-temporal reasoning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11287–11297

  67. [72]

    Lingyue Guo, Zeyu Gao, Jinye Qu, Suiwu Zheng, Runhao Jiang, Yanfeng Lu, and Hong Qiao. 2023. Transformer-based spiking neural networks for multimodal audiovisual classification.IEEE Transactions on Cognitive and Developmental Systems16, 3 (2023), 1077–1086

  68. [73]

    Zhanqiang Guo, Jiamin Wu, Yonghao Song, Jiahui Bu, Weijian Mai, Qihao Zheng, Wanli Ouyang, and Chunfeng Song. 2025. Neuro-3D: Towards 3D visual decoding from EEG signals. InProceedings of the Computer Vision and Pattern Recognition Conference. 23870–23880

  69. [74]

    Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. 2023. Mastering diverse domains through world models.arXiv preprint arXiv:2301.04104(2023)

  70. [75]

    Torkel Hafting, Marianne Fyhn, Sturla Molden, May-Britt Moser, and Edvard I Moser. 2005. Microstructure of a spatial map in the entorhinal cortex.Nature436, 7052 (2005), 801–806

  71. [76]

    Michael L Hammock, Alexander Chortos, Benjamin CK Tee, Jeffrey BH Tok, and Zhenan Bao. 2013. 25th anniversary article: The evolution of electronic skin (e-skin): A brief history, design considerations, and recent progress.Advanced Materials25, 42 (2013), 5997–6038

  72. [77]

    Demis Hassabis and Eleanor A Maguire. 2007. Deconstructing episodic memory with construction.Trends in cognitive sciences11, 7 (2007), 299–306

  73. [78]

    Hananel Hazan, Daniel J Saunders, Hassaan Khan, Devdhar Patel, Darpan T Sanghavi, Hava T Siegelmann, and Robert Kozma. 2018. Bindsnet: A machine learning-oriented spiking neural networks library in python.Frontiers in neuroinformatics12 (2018), 89. Preprint (September 2025). M...

  74. [79]

    Quentin Herau, Nathan Piasco, Moussab Bennehar, Luis Roldao, Dzmitry Tsishkou, Cyrille Migniot, Pascal Vasseur, and Cédric Demonceaux. 2024. Soac: Spatio-temporal overlap-aware multi-sensor calibration using neural radiance fields. InProceedings of the IEEE/CVF Conference on C...

  75. [80]

    Nguyen, Van-Dinh Nguyen, Yong Xiao, and Eryk Dutkiewicz

    Nguyen Quang Hieu, Dinh Thai Hoang, Diep N. Nguyen, Van-Dinh Nguyen, Yong Xiao, and Eryk Dutkiewicz. 2024. Enhancing Immersion and Presence in the Metaverse With Over-the-Air Brain-Computer Interface.IEEE Transactions on Wireless Communications23, 12 (2024), 18532–18548

  76. [81]

    Nguyen Quang Hieu, Dinh Thai Hoang, Dusit Niyato, Diep N Nguyen, Dong In Kim, and Abbas Jamalipour. 2023. Joint power allocation and rate control for rate splitting multiple access networks with covert communications.IEEE Transactions on Communications71, 4 (2023), 2274–2287

  77. [82]

    Nguyen Quang Hieu, Minh Nguyen, Dinh Thai Hoang, Diep N Nguyen, and Eryk Dutkiewicz. 2024. Point Cloud Compression with Bits-back Coding.arXiv preprint arXiv:2410.18115(2024)

  78. [83]

    Stephen C Hirtle and John Jonides. 1985. Evidence of hierarchies in cognitive maps.Memory & cognition13, 3 (1985), 208–217

  79. [84]

    Tomas Hodan, Frank Michel, Eric Brachmann, Wadim Kehl, Anders GlentBuch, Dirk Kraft, Bertram Drost, Joel Vidal, Stephan Ihrke, Xenophon Zabulis, et al. 2018. Bop: Benchmark for 6d object pose estimation. InProceedings of the European conference on computer vision (ECCV). 19–34

  80. [85]

    Newton Howard and Erik Cambria. 2013. Intention awareness: Improving upon situation awareness in human-centric environments.Human-centric Computing and Information Sciences3, 9 (2013)

  81. [86]

    Binghao Huang, Yixuan Wang, Xinyi Yang, Yiyue Luo, and Yunzhu Li. 2025. 3D-ViTac: Learning Fine-Grained Manipulation with Visuo-Tactile Sensing. InConference on Robot Learning. 2557–2578

  82. [87]

    Yinghao Huang, Manuel Kaufmann, Emre Aksan, Michael J Black, Otmar Hilliges, and Gerard Pons-Moll. 2018. Deep inertial poser: Learning to reconstruct human pose from sparse inertial measurements in real time.ACM Transactions on Graphics (TOG)37, 6 (2018), 1–15

  83. [88]

    Black, Otmar Hilliges, and Gerard Pons-Moll

    Yinghao Huang, Manuel Kaufmann, Emre Aksan, Michael J. Black, Otmar Hilliges, and Gerard Pons-Moll. 2018. Deep Inertial Poser Learning to Reconstruct Human Pose from SparseInertial Measurements in Real Time.ACM Transactions on Graphics, (Proc. SIGGRAPH Asia)37, 6 (Nov. 2018), ...

  84. [89]

    Intel. [n. d.]. Intel Movidius Myriad X VPU. Online. https://www.intel.com/content/www/us/en/products/sku/ 125926/intel-movidius-myriad-x-vision-processing-unit-4gb/specifications.html

  85. [90]

    Lucia F Jacobs and Françoise Schenk. 2003. Unpacking the cognitive map: the parallel map theory of hippocampal function.Psychological review110, 2 (2003), 285

  86. [91]

    Ayush Jain, Nikolaos Gkanatsios, Ishita Mediratta, and Katerina Fragkiadaki. 2022. Bottom up top down detection transformers for language grounding in images and point clouds. InEuropean Conference on Computer Vision. Springer, 417–433

  87. [92]

    Bo Jiang, Yao Lu, Guangming Lu, and Bob Zhang. 2024. QFormer: An efficient quaternion transformer for image denoising. InProc. 33rd Int. Joint Conf. Artif. Intell. 4237–4245

  88. [93]

    Pengcheng Jiang, Lang Cao, Cao Danica Xiao, Parminder Bhatia, et al. 2024. Kg-fit: Knowledge graph fine-tuning upon open-world knowledge.Advances in Neural Information Processing Systems37 (2024), 136220–136258

  89. [94]

    Tian Jin, Gheorghe-Teodor Bercea, Tung D Le, Tong Chen, Gong Su, Haruki Imai, et al. 2020. Compiling onnx neural network models using mlir.arXiv preprint arXiv:2008.08272(2020)

  90. [95]

    Roland S Johansson and J Randall Flanagan. 2009. Coding and use of tactile signals from the fingertips in object manipulation tasks.Nature Reviews Neuroscience10, 5 (2009), 345–359

  91. [96]

    Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. 2017. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. InProceedings of the IEEE conference on computer vision and pattern recog...

  92. [97]

    Kemp, Aaron Edsinger, and Eduardo R

    Charles C. Kemp, Aaron Edsinger, and Eduardo R. Torres-Jara. 2007. Challenges for robot manipulation in human environments.IEEE Robotics & Automation Magazine14, 1 (2007), 20–29. doi:10.1109/MRA.2007.339605

  93. [98]

    Amirhossein Khalilian-Gourtani, Ran Wang, Xupeng Chen, Leyao Yu, Patricia Dugan, Daniel Friedman, Werner Doyle, Orrin Devinsky, Yao Wang, and Adeen Flinker. 2024. A corollary discharge circuit in human speech.Proceedings of the National Academy of Sciences121, 50 (2024), e2404121121

  94. [99]

    Sungmin Kim, Cecilia Laschi, and Barry Trimmer. 2016. Soft skin for tactile sensing.Advanced Materials28, 22 (2016), 4560–4572. doi:10.1002/adma.201505115

  95. [100]

    Diederik Kingma, Tim Salimans, Ben Poole, and Jonathan Ho. 2021. Variational diffusion models.Advances in neural information processing systems34 (2021), 21696–21707

  96. [101]

    David C Knill and Alexandre Pouget. 2004. The Bayesian brain: the role of uncertainty in neural coding and computation.TRENDS in Neurosciences27, 12 (2004), 712–719

  97. [102]

    Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, et al. 2017. Ai2-thor: An interactive 3d environment for visual ai.arXiv preprint arXiv:1712.05474(2017). Preprint (September 2025). 50 Manh, et al

  98. [103]

    Anton Komarichev et al. 2022. Polyglot: Learning a Coordinated Semantic-Geometric Latent Space for Multimodal Agents.Computer Aided Geometric Design(2022)

  99. [104]

    Benjamin Kuipers. 1978. Modeling spatial knowledge.Cognitive science2, 2 (1978), 129–153

  100. [105]

    Yuzhi Lai, Shenghai Yuan, Boya Zhang, Benjamin Kiefer, Peizheng Li, and Andreas Zell. 2025. FAM-HRI: Foundation- Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech.arXiv preprint arXiv:2503.16492 (2025)

  101. [106]

    Mike Lambeta, Shuwen Chou, Rui Tian, et al . 2020. DIGIT: A novel compact high-resolution tactile sensor for dexterous manipulation. InRobotics: Science and Systems (RSS)

  102. [107]

    Kun Lan, Haoran Li, Haolin Shi, Wenjun Wu, Lin Wang, and Yong Liao. 2024. 2d-guided 3d gaussian segmentation. In2024 Asian Conference on Communication and Networks (ASIANComNet). IEEE, 1–5

  103. [108]

    Yann LeCun. 2022. A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27.Open Review62, 1 (2022), 1–62

  104. [109]

    Lee, Jihyeon Je, Chanho Park, Mikaela Angelina Uy, Leonidas Guibas, and Minhyuk Sung

    Phillip Y. Lee, Jihyeon Je, Chanho Park, Mikaela Angelina Uy, Leonidas Guibas, and Minhyuk Sung. 2025. Perspective- Aware Reasoning in Vision-Language Models via Mental Imagery Simulation. InCVPR

  105. [110]

    Stefan Leutgeb, Jill K Leutgeb, Carol A Barnes, Edvard I Moser, et al. 2005. Independent codes for spatial and episodic memory in hippocampal neuronal ensembles.Science309, 5734 (2005), 619–623

  106. [111]

    Chengzu Li, Wenshan Wu, Huanyu Zhang, Yan Xia, Shaoguang Mao, Li Dong, et al. 2025. Imagine While Reasoning in Space: Multimodal Visualization-of-Thought. InForty-second International Conference on Machine Learning

  107. [112]

    Chengshu Li, Fei Xia, Roberto Martín-Martín, Michael Lingelbach, Sanjana Srivastava, Bokui Shen, Kent Elliott Vainio, Cem Gokmen, Gokul Dharan, Tanish Jain, et al. 2022. iGibson 2.0: Object-Centric Simulation for Robot Learning of Everyday Household Tasks. InConference on Robo...

  108. [113]

    Haoran Li, Haolin Shi, Wenli Zhang, Wenjun Wu, Yong Liao, Lin Wang, Lik-hang Lee, and Peng Yuan Zhou. 2024. Dreamscene: 3d gaussian-based text-to-3d scene generation via formation pattern sampling. InEuropean Conference on Computer Vision. Springer, 214–230

  109. [114]

    Jiangmeng Li, Yifan Jin, Hang Gao, Wenwen Qiang, Changwen Zheng, and Fuchun Sun. 2024. Hierarchical topology isomorphism expertise embedded graph contrastive learning. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 13518–13527

  110. [115]

    Jinzhou Li, Tianhao Wu, Jiyao Zhang, et al. 2025. Adaptive Visuo-Tactile Fusion with Predictive Force Attention for Dexterous Manipulation.arXiv preprint arXiv:2505.13982(2025)

  111. [116]

    Shengyu Li, Shuolong Chen, et al. 2025. Accurate and automatic spatiotemporal calibration for multi-modal sensor system based on continuous-time optimization.Information Fusion120 (2025), 103071

  112. [117]

    Sicheng Li, Hao Li, Yiyi Liao, and Lu Yu. 2024. Nerfcodec: Neural feature compression meets neural radiance fields for memory-efficient scene representation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 21274–21283

  113. [118]

    Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chonghao Sima, Tong Lu, Qiao Yu, and Jifeng Dai. 2024. Bevformer: learning bird’s-eye-view representation from lidar-camera via spatiotemporal transformers.IEEE Transactions on Pattern Analysis and Machine Intelligence(2024)

  114. [119]

    Zhong-Zhi Li, Duzhen Zhang, Ming-Liang Zhang, Jiaxin Zhang, Zengyan Liu, Yuxuan Yao, Haotian Xu, Junhao Zheng, Pei-Jie Wang, Xiuyi Chen, et al. 2025. From system 1 to system 2: A survey of reasoning large language models.arXiv preprint arXiv:2502.17419(2025)

  115. [120]

    Wei Yang Bryan Lim, Nguyen Cong Luong, Dinh Thai Hoang, Yutao Jiao, Ying-Chang Liang, Qiang Yang, Dusit Niyato, and Chunyan Miao. 2020. Federated learning in mobile edge networks: A comprehensive survey.IEEE communications surveys & tutorials22, 3 (2020), 2031–2063

  116. [121]

    Changyi Lin, Han Zhang, et al. 2023. 9dtact: A compact vision-based tactile sensor for accurate 3d shape reconstruction and generalizable 6d force estimation.IEEE Robotics and Automation Letters9, 2 (2023), 923–930

  117. [122]

    Wang Lin and Wang Liang. 2017. Neural Circuit Basis of Cognitive Map.SCI, CA, Scopus, AJ44, 3 (2017), 187

  118. [123]

    Fangchen Liu, Chuanyu Li, Yihua Qin, Ankit Shaw, Jing Xu, Pieter Abbeel, and Rui Chen. 2025. Vitamin: Learning contact-rich tasks through robot-free visuo-tactile manipulation interface.arXiv preprint arXiv:2504.06156(2025)

  119. [124]

    Hao Liu, Runguo Wei, Geng Tu, Jiali Lin, Dazhi Jiang, and Erik Cambria. 2025. Knowing What and Why: Causal Emotion Entailment for Emotion Recognition in Conversations.Expert Systems with Applications274 (2025), 126924

  120. [126]

    Jian Liu, Wei Sun, et al . 2023. Robotic Continuous Grasping System by Shape Transformer-Guided Multiobject Category-Level 6-D Pose Estimation.IEEE Transactions on Industrial Informatics19, 11 (2023), 11171–11181

  121. [127]

    Jian Liu, Wei Sun, Hui Yang, Pengchao Deng, Chongpei Liu, et al. 2025. Diff9d: Diffusion-based domain-generalized category-level 9-dof object pose estimation.IEEE Transactions on Pattern Analysis and Machine Intelligence(2025). Preprint (September 2025). Mind Meets Space: Reth...

  122. [128]

    Jian Liu, Wei Sun, Hui Yang, Zhiwen Zeng, Chongpei Liu, et al. 2024. Deep Learning-Based Object Pose Estimation: A Comprehensive Survey.arXiv preprint arXiv:2405.07801(2024)

  123. [130]

    Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. 2015. SMPL: a skinned multi-person linear model.ACM Transactions on Graphics (TOG)34, 6 (2015), 1–16

  124. [131]

    Haodong Lu, Xinyu Zhang, Kristen Moore, Jason Xue, et al . 2025. Continual Learning on CLIP via Incremental Prompt Tuning with Intrinsic Textual Anchors.arXiv preprint arXiv:2505.20680(2025)

  125. [132]

    Chenyang Ma, Kai Lu, Ta-Ying Cheng, Niki Trigoni, and Andrew Markham. 2024. SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors. InAdvances in Neural Information Processing Systems (NeurIPS)

  126. [133]

    Wufei Ma, Haoyu Chen, Guofeng Zhang, Yu-Cheng Chou, Celso M de Melo, and Alan Yuille. 2024. 3dsrbench: A comprehensive 3d spatial reasoning benchmark.arXiv preprint arXiv:2412.07825(2024)

  127. [134]

    Naureen Mahmood, Nima Ghorbani, et al. 2019. AMASS: Archive of motion capture as surface shapes. InProceedings of the IEEE/CVF international conference on computer vision. 5442–5451

  128. [135]

    Rui Mao, Guanyi Chen, Xiao Li, Mengshi Ge, and Erik Cambria. 2025. A Comparative Analysis of Metaphorical Cognition in ChatGPT and Human Minds.Cognitive Computation17 (2025), 35

  129. [136]

    Rui Mao, Guanyi Chen, Xulang Zhang, Frank Guerin, and Erik Cambria. 2024. GPTEval: A Survey on Assessments of ChatGPT and GPT-4. InLREC-COLING. 7844–7866

  130. [137]

    Rui Mao, Qian Liu, Xiao Li, Erik Cambria, and Amir Hussain. 2025. Bridging Minds and Machines: Toward an Integration of AI and Cognitive Science.arXiv preprint arXiv:2508.20674(2025)

  131. [138]

    Tobias Meilinger et al. 2016. Qualitative differences in memory for vista and environmental spaces are caused by opaque borders, not movement or successive presentation.Cognition155 (2016), 77–95

  132. [139]

    Meta Platforms, Inc. 2025. Codec Avatars: Meta Immersive Telepresence. https://www.meta.com/emerging-tech/ codec-avatars/. Accessed: 2025-08-21

  133. [140]

    Meta Platforms, Inc. 2025. Meta Quest Virtual Reality Headset. https://www.meta.com/quest/. Accessed: 2025-08-21

  134. [141]

    John C Middlebrooks and David M Green. 1991. Sound localization by human listeners.Annual Review of Psychology 42 (1991), 135–159. doi:10.1146/annurev.ps.42.020191.001031

  135. [142]

    Edvard I Moser, Emilio Kropff, and May-Britt Moser. 2008. Place cells, grid cells, and the brain’s spatial representation system.Annu. Rev. Neurosci.31, 1 (2008), 69–89

  136. [143]

    NVIDIA Corporation. 2023. Jetson Orin Nano Series. Online. https://www.nvidia.com/en-us/autonomous-machines/ embedded-systems/

  137. [144]

    NVIDIA Corporation. 2025. NVIDIA Aerial Expands With New Tools for Building AI-Native Wireless Networks. Online. https://blogs.nvidia.com/blog/aerial-telecom-research/

  138. [146]

    John O’Keefe and Jonathan Dostrovsky. 1971. The hippocampus as a spatial map: preliminary evidence from unit activity in the freely-moving rat.Brain research(1971)

  139. [147]

    J O’Keefe and L Nadel. 1978. The Hippocampus as a Cognitive Map

  140. [148]

    Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748(2018)

  141. [149]

    OpenAI. 2025. Introducing ChatGPT agent: bridging research and action. https://openai.com/index/introducing- chatgpt-agent/ Accessed: July. 30, 2025

  142. [150]

    Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. 2023. Dinov2: Learning robust visual features without supervision.arXiv preprint arXiv:2304.07193(2023)

  143. [151]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems35 ...

  144. [152]

    2013.Imagery and verbal processes

    Allan Paivio. 2013.Imagery and verbal processes. Psychology Press

  145. [153]

    Kramer, Pascal Bérard, and Robert J

    Yong-Lae Park, Carmel Majidi, Rebecca K. Kramer, Pascal Bérard, and Robert J. Wood. 2010. Hyperelastic pressure sensing with a liquid-embedded elastomer.Advanced Functional Materials20, 21 (2010), 3547–3555

  146. [154]

    Alexander Pashevich, Cordelia Schmid, and Chen Sun. 2021. Episodic transformer for vision-and-language navigation. InProceedings of the IEEE/CVF International Conference on Computer Vision. 15942–15952

  147. [155]

    Eva Zita Patai and Hugo J Spiers. 2021. The versatile wayfinder: prefrontal contributions to spatial navigation.Trends in cognitive sciences25, 6 (2021), 520–533. Preprint (September 2025). 52 Manh, et al

  148. [156]

    Matthew J Pearson, Michael H Evans, and Robert Trew. 2007. A biomimetic whisker for robots. InProceedings of the IEEE/ASME International Conference on Advanced Intelligent Mechatronics. IEEE, 1–6

  149. [157]

    Roberto Pegurri, Eugenio Moro, Francesco Linsalata, Jakob Hoydis, and Umberto Spagnolini. 2025. Towards Digital Network Twins: Full-Stack and Multi-Stack Solution for 6G Simulations. In2025 IEEE Wireless Communications and Networking Conference (WCNC). IEEE, 1–3

  150. [158]

    Christian-Gernot Pehle and Jens Egholm Pedersen. 2021. Norse-A deep learning library for spiking neural networks. Zenodo(2021)

  151. [159]

    Will D Penny, Peter Zeidman, and Neil Burgess. 2013. Forward and backward inference in spatial cognition.PLoS computational biology9, 12 (2013), e1003383

  152. [160]

    Markus Pettersen, Frederik Rogge, and Mikkel Elle Lepperød. 2024. Learning Place Cell Representations and Context-Dependent Remapping. InAdvances in Neural Information Processing Systems, Vol. 37. 244–269

  153. [161]

    Giovanni Pezzulo, Francesco Rigoli, and Karl J Friston. 2018. Hierarchical active inference: a theory of motivated control.Trends in cognitive sciences22, 4 (2018), 294–306

  154. [162]

    Brien M Posey. 2022. What is the akida event domain neural processor?2020(2022)

  155. [163]

    Xavier Puig, Eric Undersander, Andrew Szot, Mikael Dallaire Cote, Tsung-Yen Yang, Ruslan Partsey, Ruta Desai, Alexander William Clegg, Michal Hlavac, So Yeon Min, et al. 2023. Habitat 3.0: A co-habitat for humans, avatars and robots.arXiv preprint arXiv:2310.13724(2023)

  156. [164]

    Xinyuan Qian, Jiaran Gao, Yaodan Zhang, Qiquan Zhang, Hexin Liu, Leibny Paola Garcia, and Haizhou Li. 2025. SAV-SE: Scene-aware Audio-Visual Speech Enhancement with Selective State Space Model.IEEE Journal of Selected Topics in Signal Processing(2025)

  157. [165]

    Minghan Qin, Wanhua Li, Jiawei Zhou, Haoqian Wang, and Hanspeter Pfister. 2024. Langsplat: 3d language gaussian splatting. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 20051–20060

  158. [167]

    Rajesh PN Rao and Dana H Ballard. 1999. Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects.Nature neuroscience2, 1 (1999), 79–87

  159. [168]

    Aditya Ravi, Sudeep Manandhar, Kasi Viswanathan, Nemanja Djuric, Sarthak Sharma, and Ibrahim Pirk

  160. [169]

    Sahithya Ravi, Gabriel Sarch, Vibhav Vineet, et al . 2025. Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames.arXiv preprint arXiv:2505.24257(2025)

  161. [170]

    Edmund T Rolls. 2013. The mechanisms for pattern completion and pattern separation in the hippocampus.Frontiers in systems neuroscience7 (2013), 74

  162. [171]

    arXiv:2404.09995 [cs.CV]

    OutSight: A Multi-Sensor 3D Object Detection Framework with Cross-Modal Representation Learning. arXiv:2404.09995 [cs.CV]

  163. [172]

    Rylan Schaeffer et al. 2022. No free lunch from deep learning in neuroscience: A case study through models of the entorhinal-hippocampal circuit.Advances in neural information processing systems35 (2022), 16052–16067

  164. [173]

    Rylan Schaeffer, Mikail Khona, Tzuhsuan Ma, Cristobal Eyzaguirre, Sanmi Koyejo, and Ila Fiete. 2023. Self-Supervised Learning of Representations for Space Generates Multi-Modular Grid Cells. InAdvances in Neural Information Processing Systems, Vol. 36. 23140–23157

  165. [174]

    Sasha Salter, Richard Warren, Collin Schlager, Adrian Spurr, Shangchen Han, Rohin Bhasin, Yujun Cai, Peter Walk- ington, Anuoluwapo Bolarinwa, Robert J Wang, et al. 2024. emg2pose: A large and diverse benchmark for surface electromyographic hand pose estimation.Advances in Neu...

  166. [175]

    Amir-Hossein Shahidzadeh, Gabriele Caddeo, Koushik Alapati, Lorenzo Natale, Cornelia Fermüller, and Yiannis Aloimonos. 2024. FeelAnyForce: Estimating contact force feedback from tactile sensation for vision-based tactile sensors.arXiv preprint arXiv:2410.02048(2024)

  167. [176]

    Chen Shang, Jiadong Yu, and Dinh Thai Hoang. 2025. Energy-efficient and intelligent ISAC in V2X networks with spiking neural networks-driven DRL.arXiv preprint arXiv:2501.01038(2025)

  168. [177]

    Paul Hongsuck Seo et al. 2023. Avformer: Injecting vision into frozen speech models for zero-shot av-asr. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 22922–22931

  169. [178]

    Xinyu Shi, Zecheng Hao, and Zhaofei Yu. 2024. Spikingresformer: Bridging resnet and vision transformer in spiking neural networks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 5610–5619

  170. [179]

    Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim. 2017. Continual learning with deep generative replay. Advances in neural information processing systems30 (2017)

  171. [180]

    Qiongjie Shi, Beibei Dong, Tian-Li He, Yunlong Zi, Zhihong Zeng, et al. 2020. Electronic skin: Advances and future prospects.Science370, 6518 (2020), 998–1004

  172. [181]

    Kirk Shung and Mark Zippuro

    K. Kirk Shung and Mark Zippuro. 1996. Ultrasonic transducers and arrays.IEEE Engineering in Medicine and Biology Magazine15, 6 (1996), 20–30

  173. [182]

    Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. 2012. Indoor segmentation and support inference from rgbd images. InEuropean conference on computer vision. Springer, 746–760

  174. [183]

    Jonathan P Shine, José P Valdés-Herrera, Mary Hegarty, and Thomas Wolbers. 2016. The human retrosplenial cortex and thalamus code head direction in a global reference frame.Journal of Neuroscience36, 24 (2016), 6371–6381. Preprint (September 2025). Mind Meets Space: Rethinking...

  175. [184]

    Daniel Sliwowski, Shail Jadav, Sergej Stanovcic, Jedrzej Orbik, et al. 2025. Reassemble: A multimodal dataset for contact-rich robotic assembly and disassembly.arXiv preprint arXiv:2502.05086(2025)

  176. [185]

    Ben Sorscher, Gabriel Mel, Surya Ganguli, and Samuel Ocko. 2019. A unified theory for the origin of grid cells through the lens of pattern formation. InAdvances in neural information processing systems, Vol. 32

  177. [186]

    Viswanath Sivakumar, Jeffrey Seely, Alan Du, Sean Bittner, Adam Berenzweig, Anuoluwapo Bolarinwa, Alex Gramfort, and Michael Mandel. 2024. emg2qwerty: A large dataset with baselines for touch typing using surface electromyog- raphy.Advances in Neural Information Processing Sys...

  178. [187]

    John F Stein. 1992. The representation of egocentric space in the posterior parietal cortex.Behavioral and Brain Sciences15, 4 (1992), 691–700

  179. [188]

    Marcel Stimberg et al. 2019. Brian 2, an intuitive and efficient neural simulator.elife8 (2019), e47314

  180. [189]

    Kimberly L Stachenfeld, Matthew M Botvinick, and Samuel J Gershman. 2017. The hippocampus as a predictive map. Nature neuroscience20, 11 (2017), 1643–1653

  181. [190]

    Tengine. 2021. Deploy Edge AI App In a Second With Tengine and SuperEdge. Online. https://github.com/OAID/ Tengine/blob/tengine-lite/doc/docs_en/source_compile/deploy_SuperEdge.md

  182. [191]

    Tesla, Inc. 2025. Tesla Optimus Humanoid Robot. https://www.tesla.com/ai. Accessed: 2025-08-21

  183. [192]

    Andrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans, and Yiliand others Zhao. 2021. Habitat 2.0: Training home assistants to rearrange their habitat.Advances in neural information processing systems34 (2021), 251–266

  184. [193]

    1983.Elements of Episodic Memory

    Endel Tulving. 1983.Elements of Episodic Memory. Oxford University Press, Oxford, GB

  185. [194]

    Barbara Tversky. 1981. Distortions in memory for maps.Cognitive psychology13, 3 (1981), 407–433

  186. [195]

    Edward C Tolman. 1948. Cognitive maps in rats and men.Psychological review55, 4 (1948), 189

  187. [196]

    Jiang Wang, Yaozhong Kang, Linya Fu, et al. 2025. Observability-Aware Active Calibration of Multi-Sensor Extrinsics for Ground Robots via Online Trajectory Optimization.arXiv preprint arXiv:2506.13420(2025)

  188. [197]

    Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, et al. 2024. A survey on large language model based autonomous agents.Frontiers of Computer Science18, 6 (2024), 186345

  189. [198]

    Ungerleider and Mortimer Mishkin

    Leslie G. Ungerleider and Mortimer Mishkin. 1982. Two cortical visual systems. InAnalysis of Visual Behavior, D. J. Ingle, M. A. Goodale, and R. J. W. Mansfield (Eds.). MIT Press, 549–586

  190. [199]

    Wenqi Wang, Reuben Tan, Pengyue Zhu, Jianwei Yang, Zhengyuan Yang, Lijuan Wang, et al. 2025. SITE: towards Spatial Intelligence Thorough Evaluation.arXiv preprint arXiv:2505.05456(2025)

  191. [200]

    Zirui Wang, Xinran Zhao, et al. 2024. HiMemFormer: Hierarchical Memory-Aware Transformer for Multi-Agent Action Anticipation.arXiv preprint arXiv:2411.01455(2024)

  192. [201]

    Lin Wang and Kuk-Jin Yoon. 2021. Knowledge distillation and student-teacher learning for visual intelligence: A review and new outlooks.IEEE transactions on pattern analysis and machine intelligence44, 6 (2021), 3048–3068

  193. [202]

    James CR Whittington, David McCaffary, Jacob JW Bakermans, and Timothy EJ Behrens. 2022. How to build a cognitive map.Nature neuroscience25, 10 (2022), 1257–1272

  194. [203]

    James CR Whittington, Timothy H Muller, Shirley Mark, Guifen Chen, Caswell Barry, Neil Burgess, and Timothy EJ Behrens. 2020. The Tolman-Eichenbaum machine: unifying space and relational memory through generalization in the hippocampal formation.Cell183, 5 (2020), 1249–1263

  195. [204]

    Greg Wayne, Chia-Chun Hung, David Amos, Mehdi Mirza, Arun Ahuja, Agnieszka Grabska-Barwinska, Jack Rae, Piotr Mirowski, Joel Z Leibo, Adam Santoro, et al. 2018. Unsupervised predictive memory in a goal-directed agent. arXiv preprint arXiv:1803.10760(2018)

  196. [205]

    Christopher Widdowson and Ranxiao Frances Wang. 2022. Human navigation in curved spaces.Cognition218 (2022), 104923

  197. [206]

    Junfei Wu, Jian Guan, Kaituo Feng, Qiang Liu, Shu Wu, Liang Wang, et al. 2025. Reinforcing spatial reasoning in vision-language models with interwoven thinking and visual drawing.arXiv preprint arXiv:2506.09965(2025)

  198. [207]

    James CR Whittington, Joseph Warren, and Timothy EJ Behrens. 2021. Relating transformers to models and neural representations of the hippocampal formation.arXiv preprint arXiv:2112.04035(2021)

  199. [208]

    Eric Xing, Mingkai Deng, et al. 2025. Critiques of World Models.arXiv preprint arXiv:2507.05169(2025)

  200. [209]

    Fangzhi Xu, Qika Lin, Jiawei Han, et al . 2025. Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond.Transactions on Knowledge and Data Engineering37 (2025), 1620–1634

  201. [210]

    Yu Xiang, Wonhui Kim, Wei Chen, Jingwei Ji, Christopher Choy, Hao Su, et al. 2016. Objectnet3d: A large scale database for 3d object recognition. InEuropean conference on computer vision. Springer, 160–176

  202. [211]

    Fengyu Yang, Chao Feng, Ziyang Chen, Hyoungseob Park, Daniel Wang, Yiming Dou, Ziyao Zeng, Xien Chen, Rit Gangopadhyay, Andrew Owens, et al. 2024. Binding touch to everything: Learning unified multimodal tactile representations. InProceedings of the IEEE/CVF Conference on Comp...

  203. [212]

    Fanbo Yang, Ankur Handa, Vladlen Koltun, and Thomas Funkhouser. 2024. 3D-Mem: 3D Scene Memory for Embodied Reasoning.arXiv preprint arXiv:2403.02395(2024)

  204. [213]

    Ziang Yan, Zhilin Li, Yinan He, Chenting Wang, Kunchang Li, Xinhao Li, Xiangyu Zeng, Zilei Wang, Yali Wang, Yu Qiao, et al. 2025. Task preference optimization: Improving multimodal large language models with vision task Preprint (September 2025). 54 Manh, et al. alignment. InP...

  205. [214]

    Man Yao, Guangshe Zhao, Hengyu Zhang, Yifan Hu, Lei Deng, Yonghong Tian, Bo Xu, and Guoqi Li. 2023. Attention spiking neural networks.IEEE transactions on pattern analysis and machine intelligence45, 8 (2023), 9393–9410

  206. [215]

    Baiqiao Yin, Qineng Wang, Pingyue Zhang, Jianshu Zhang, Kangrui Wang, Zihan Wang, et al. 2025. Spatial Mental Modeling from Limited Views.arXiv preprint arXiv:2506.21458(2025)

  207. [216]

    Jihan Yang, Shusheng Yang, Anjali W Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie. 2025. Thinking in space: How multimodal large language models see, remember, and recall spaces. InProceedings of the Computer Vision and Pattern Recognition Conference. 10632–10643

  208. [217]

    Andy Zeng, Shuran Song, Matthias Nießner, Matthew Fisher, Jianxiong Xiao, and Thomas Funkhouser. 2017. 3dmatch: Learning local geometric descriptors from rgb-d reconstructions. InProceedings of the IEEE conference on computer vision and pattern recognition. 1802–1811

  209. [218]

    Jingyi Zhang, Jiaxing Huang, Sheng Jin, and Shijian Lu. 2024. Vision-language models for vision tasks: A survey. IEEE transactions on pattern analysis and machine intelligence46, 8 (2024), 5625–5644

  210. [219]

    Vladimir Yugay, Yue Li, Theo Gevers, and Martin R. Oswald. 2023. Gaussian-SLAM: Photo-realistic Dense SLAM with Gaussian Splatting. arXiv:2312.10070 [cs.CV]

  211. [220]

    Ruichen Zhang, Shunpu Tang, Yinqiu Liu, Dusit Niyato, et al . 2025. Toward agentic ai: generative information retrieval inspired intelligent communications and networking.arXiv preprint arXiv:2502.16866(2025)

  212. [221]

    Siwei Zhang, Qianli Ma, Yan Zhang, Zhiyin Qian, Taein Kwon, Marc Pollefeys, Federica Bogo, and Siyu Tang. 2022. Egobody: Human body shape and motion of interacting people from head-mounted devices. InEuropean conference on computer vision. Springer, 180–200

  213. [222]

    Qiyang Zhang, Xiang Li, Xiangying Che, Xiao Ma, Ao Zhou, Mengwei Xu, et al. 2022. A comprehensive benchmark of deep learning libraries on mobile devices. InProceedings of the ACM Web Conference 2022. 3298–3307

  214. [223]

    Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip HS Torr, and Vladlen Koltun. 2021. Point transformer. InProceedings of the IEEE/CVF international conference on computer vision. 16259–16268

  215. [224]

    Jialiang Zhao, Yuxiang Ma, Lirui Wang, and Edward H Adelson. 2024. Transferable tactile transformers for represen- tation learning across diverse sensors and tasks.arXiv preprint arXiv:2406.13640(2024)

  216. [225]

    Zhaoliang Zhang, Tianchen Song, Yongjae Lee, Li Yang, Cheng Peng, Rama Chellappa, and Deliang Fan. 2024. Lp-3dgs: Learning to prune 3d gaussian splatting.Advances in Neural Information Processing Systems37 (2024), 122434–122457

  217. [226]

    Gaoyue Zhou, Hengkai Pan, Yann LeCun, and Lerrel Pinto. 2024. Dino-wm: World models on pre-trained visual features enable zero-shot planning.arXiv preprint arXiv:2411.04983(2024)

  218. [227]

    Alex Zihao Zhu, Dinesh Thakur, Tolga Özaslan, et al. 2018. The multivehicle stereo event camera dataset: An event camera dataset for 3D perception.IEEE Robotics and Automation Letters3, 3 (2018), 2032–2039

  219. [228]

    Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang, Xin Yan, Yilun Du, Yining Hong, and Chuang Gan. 2024. 3d-vla: A 3d vision-language-action generative world model.arXiv preprint arXiv:2403.09631(2024)

  220. [229]

    Howe Yuan Zhu, Nguyen Quang Hieu, Dinh Thai Hoang, et al . 2024. A Human-Centric Metaverse Enabled by Brain-Computer Interface: A Survey.IEEE Communications Surveys & Tutorials26, 3 (2024), 2120–2145. Preprint (September 2025)

  221. [231]

    Deyao Zhu, Jun Chen, et al. 2024. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. InThe Twelfth International Conference on Learning Representations

  222. [2024]

    InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 19383–19400

  223. [2025]

    InProceedings of the Computer Vision and Pattern Recognition Conference

    Perception tokens enhance visual reasoning in multimodal language models. InProceedings of the Computer Vision and Pattern Recognition Conference. 3836–3845

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.