Pith. sign in

REVIEW 4 major objections 4 minor 2 cited by

Nature's Insight: A Novel Framework and Comprehensive Analysis of Agentic Reasoning Through the Lens of Neuroscience

T0 review · 4 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This paper proposes a neuroscience-inspired taxonomy that sorts agentic reasoning into perceptual, dimensional, logical, and interactive types, and maps each to specific brain subsystems and AI agent modules.

desk verdict Useful survey of agentic reasoning wrapped in a neuroscience framework whose one-to-one brain-module mapping does not hold up; the compilation is worth peer review, the framework needs heavy reframing. read the letter →

arxiv 2505.05515 v1 pith:B2YPLKF4 submitted 2025-05-07 q-bio.NC cs.LG

classification q-bio.NCcs.LG
keywords agenticreasoningcognitiveneuroscienceneuroscience-inspiredAItaxonomyperceptualdimensionallogicalinteractive
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to give agentic reasoning in AI a single organizing account drawn from cognitive neuroscience: reasoning is a hybrid, recursive, multistep process that starts with multimodal perception and ends in action, and every AI reasoning method can be placed somewhere inside that arc. Its central proposal is a four-part taxonomy—perceptual, dimensional, logical, and interactive reasoning—with each type inspired by distinct functional subsystems in the human brain and paired with a corresponding module in an AI agent architecture. If this framework is right, researchers gain a shared language for comparing methods, aligning benchmarks, and designing agents whose reasoning is cognitively aligned rather than task-specialized. The paper demonstrates the taxonomy by classifying a broad set of methods, datasets, and applications, and by deriving new neural-inspired design directions.

What carries the argument

The load-bearing object is the paired diagram of the human reasoning brain and the agent reasoning architecture: a one-to-one mapping in which sensory cortices correspond to the multimodal input module, association areas to the information processing module, hippocampus and cortical memory to the knowledge base, and prefrontal and parietal executive circuits to the reasoning module, with the foundation model playing a dual role as understanding engine and reasoning assistant. This correspondence does the work because it converts the taxonomy into a structural claim: any reasoning method can be located by which of the four types it instantiates and by how well it fits the five-module pipeline. The mathematical formalisms—Bayes' rule, prediction error, variational free energy, and Bayesian policy optimization—supply the update dynamics, while named cognitive architectures such as ACT-R (a chunk-and-production cognitive architecture), SOAR (a symbolic production-rule architecture), and Global Workspace Theory (competition for a broadcast workspace) supply the comparison baselines from which the framework distinguishes itself.

What would settle it

Construct four task batteries that isolate perceptual, dimensional, logical, and interactive reasoning, and test models with a single module ablated (frozen knowledge base, disabled reasoning module, or no multimodal fusion). If losing a module impairs all four types equally, or if humans with focal lesions in the brain regions named for each type show no corresponding selective deficit, the proposed one-to-one mapping fails. A weaker but still useful test: if performance across the four batteries is explained entirely by task difficulty or model scale, with no selective dissociations, the taxonomy carries no predictive content.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central claim is that agentic reasoning is not a collection of tricks but a full biological-to-computational pipeline: from multimodal sensory input, through information processing into a shared representation, retrieval from a dual knowledge base, and foundation-model-assisted inference, to action and feedback that updates the system. It defines reasoning through three properties—hybrid (prior knowledge plus new information), recursive (outputs feed back as inputs), and multistep (structured progression)—and formalizes these with Bayesian updating, predictive coding, and free-energy minimization. On that foundation it identifies four core reasoning types: perceptual, tied to occipital and parietal sensory-integration circuits; dimensional, tied to parietal, prefrontal, and medial-temporal circuits for space, time, and hierarchy; logical, tied to prefrontal rule-based inference; and interactive, tied to social and fronto-parietal coordination circuits. The paper then uses this framework to reclassify existing AI methods, align benchmark datasets with reasoning types, survey embodied and virtual applications, and propose future architectures such as dynamic multimodal mixture-of-experts, dual memory systems, and neural-ODE-based continuous spatiotemporal reasoning. The contribution is the framework itself: a structured definition of agentic reasoning that spans perception to action and gives every method a place.

Load-bearing premise

The load-bearing premise is that the four reasoning types correspond to distinct functional brain subsystems and that this correspondence can be carried over one-to-one into AI agent modules; if that mapping is not real, the taxonomy is an arbitrary labeling scheme rather than a neuroscience-grounded structure.

Editorial extensions

If this is right

  • Existing AI reasoning methods can be compared on a common grid: which of the four reasoning types they instantiate, and how completely they realize the five-module perception-to-action pipeline.
  • Benchmarks become classifiable by reasoning type, so coverage gaps—such as the relative scarcity of interactive and dimensional tasks—become visible and fixable.
  • Future agent design inherits concrete architectural recommendations: selective multimodal perception, unified cross-modal representations, dual offline and online knowledge, and a foundation model used as an understanding engine plus reasoning assistant.
  • Cognitive models such as multistep prefrontal control, working-memory buffers, predictive coding, and global broadcasting become a source of new prompting and architecture ideas, just as ACT-R inspired chain-of-thought prompting.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the one-to-one brain-module mapping is an analogy until tested; a decisive extension would be a lesion-style experiment in which ablating the knowledge base or reasoning module produces the specific, dissociable deficits predicted for perceptual versus logical tasks.
  • Editorial inference: if the taxonomy is right, human neuropsychological dissociations should have analogues in AI models—for example, models with strong logical reasoning but weak interactive reasoning, mirroring patients with selective social-cognition deficits. A probe battery that looks for such double dissociations would test the framework's structure rather than its labels.
  • Editorial inference: the four categories are not claimed to be exhaustive or mutually exclusive; the paper offers examples that mix types. A natural extension is to formalize mixed-type reasoning as composition over the four primitives, turning the taxonomy into a generative grammar for reasoning tasks.
  • Editorial inference: the paper's classification of methods is qualitative; a quantitative next step would be to measure how much variance in model performance across benchmarks is explained by the four-type labels, compared with task difficulty or model scale.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes a neuroscience-inspired framework for agentic reasoning in AI. It introduces three conceptual definitions of reasoning (hybrid, recursive, multistep), reviews textbook mathematical formalisms (Bayesian inference, predictive coding, free energy, Bayesian RL), and then identifies four reasoning types—perceptual, dimensional, logical, and interactive—which it claims reflect distinct functional subsystems of the human brain. These types are mapped one-to-one onto an agent architecture consisting of multimodal input, information processing, knowledge base, foundation model, and reasoning module. The paper applies this taxonomy to survey AI reasoning methods, benchmarks, and applications, and proposes future directions including a Dynamic Multimodal Mixture-of-Experts and a dual knowledge architecture. The authors state in the introduction and conclusion that this is the first systematic examination of agentic reasoning from a neuroscience perspective, and they release an associated open-source repository.

Significance. If the proposed taxonomy and its brain-to-agent mapping were rigorously established, the framework could serve as a useful organizing principle for the rapidly growing literature on agentic reasoning, and the open-source repository would be a practical resource for the community. The paper covers a broad range of methods and benchmarks, and the mathematical equations in Section II-B are correct textbook statements. However, the central contribution is currently asserted rather than derived: the mapping between the four reasoning types and distinct neural subsystems is supported only by analogy, the mathematical foundations in Section II-B are not used to constrain or validate the taxonomy, and the paper contains an internal inconsistency between the four-type taxonomy in the text and the five-category diagram in Fig. 4. The significance is therefore conditional on substantial reframing and on either empirical or derivational support for the core mapping; as written, the framework risks being an arbitrary labeling scheme.

major comments (4)
  1. [Section II-E and Fig. 4] The taxonomy is internally inconsistent. Fig. 4's caption and diagram list five major categories of reasoning behaviors, including "relation reasoning," while Section II-E and the remainder of the paper (e.g., Section III and Fig. 5) use four categories and fold relational reasoning into perceptual reasoning. Because the four-type taxonomy is the paper's central contribution, this inconsistency must be resolved; otherwise the categorization appears imposed rather than discovered.
  2. [Section II-B, Eqs. (1)-(5)] The mathematical formulation is not connected to the proposed taxonomy. Equations (1)-(5) are standard Bayesian inference, predictive coding, free energy, and Bayesian optimization statements; nothing in the derivation identifies perceptual, dimensional, logical, and interactive reasoning as the natural decomposition of reasoning, and no equation constrains the architecture in Section II-D. The claim that the framework is "supported by mathematical foundations" is therefore not substantiated by the equations as they stand.
  3. [Sections II-D/II-E and Fig. 1] The one-to-one mapping between brain subsystems and agent modules is asserted through analogy (e.g., foundation model as memory, reasoning module as prefrontal/parietal cortex) rather than established. The authors themselves concede in Section VI that "direct biological equivalence remains unproven." Since the load-bearing claim is that the four reasoning types reflect "distinct functional subsystems" of the brain, the paper should either provide corroborating evidence (e.g., neuroimaging or lesion dissociations, or a derivation from Section II-B) or explicitly reframe the mapping as a heuristic hypothesis to be tested.
  4. [Section III and Fig. 5] The survey re-labels the four categories as perception-based, dimension-based, logic-based, and interaction-based, and assigns concrete methods to categories in ways that are not justified by the neuroscience definitions. For example, chain-of-thought methods are placed under "lingual reasoning" within perception-based reasoning, even though the text in Section III-A2 states that human reasoning does not primarily rely on language centers. This undercuts the claimed systematic alignment between the taxonomy and the classified methods and needs either a justification or a revised categorization.
minor comments (4)
  1. [Introduction and Conclusion] The claim to be "the first to systematically examine agentic reasoning from a neuroscience perspective" is a strong literature claim and should be supported with a more explicit comparison to prior surveys, or tempered to avoid overstatement.
  2. [Fig. 8] The DeepSeek R1 "game of 24" example is anecdotal and not a controlled evaluation; presenting it as evidence of a general limitation of LLM-based reasoning may mislead readers, and it should be described as an illustrative failure case or replaced with systematic results.
  3. [Table VII] There are reference inconsistencies in the benchmark table: "ReColr" should be "ReClor," and CLEVR is cited as both [197] and [238] in different places. These should be harmonized.
  4. [Section II-D] The description of the "unified electrochemical signal format" as inspiration for a modality-agnostic representation is a useful analogy, but the text should explicitly note the limits of this analogy to avoid implying a direct biophysical equivalence.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: taxonomy is asserted from external neuroscience sources, not fitted, predicted, or reduced to the paper's own inputs.

full rationale

I find no significant circularity. The paper's mathematical foundation (Sec. II-B, Eqs. 1-5) restates standard Bayesian inference, predictive coding, free energy, and Bayesian decision equations, but these equations are not used to derive the four reasoning types or the agent architecture; they are presented as background formalism. The taxonomy is introduced in Sec. II-E as a synthesis of external neuroscience literature ('Synthesizing the most widely accepted hypotheses [27], [52]'), not as a consequence of the paper's own equations. No parameter is fitted and then renamed as a prediction, and the future directions in Sec. VI are proposals rather than claimed predictions from the framework. I also find no load-bearing self-citation chain: the framework's support is drawn from external textbooks and reviews, and the novelty claim ('first to systematically examine agentic reasoning from a neuroscience perspective') is an assertion, not a mathematical result imported from the authors' prior work. The Fig. 4 vs. Sec. II-E category-count inconsistency (five vs. four reasoning types) is a correctness and internal-consistency concern, not circularity, because neither version is derived from the other by construction. Overall, the central framework is assumed rather than proven, but it is not circular in the sense of equating a prediction with its input or importing an unverified uniqueness result from self-citations.

Assumptions & free parameters 0 free parameters · 4 assumptions · 2 invented entities

The central framework relies on a set of assumptions from cognitive neuroscience and on analogies between brain modules and AI components. No free parameters are fitted, because the paper makes no quantitative claims. The invented entities are speculative future architectures with no independent evidence in the paper itself.

assumptions (4)
  • domain assumption Reasoning is a hybrid process that combines prior knowledge and new information.
    Section II-A defines reasoning as hybrid; this is a conceptual stance, not an established fact.
  • domain assumption The brain implements Bayesian inference, predictive coding, and free energy minimization.
    Section II-B relies on these theories as uncontroversial, though they are active research topics and not universally accepted.
  • ad hoc to paper Four reasoning types (perceptual, dimensional, logical, interactive) map onto distinct brain networks.
    Section II-E introduces this taxonomy and maps each type to brain regions, but the mapping is asserted, not empirically derived.
  • ad hoc to paper AI agent modules (multimodal input, information processing, knowledge base, foundation model, reasoning module) can mirror brain structures and functions.
    Section II-D proposes this analogy as the design basis for the framework, without validation.
invented entities (2)
  • Dynamic Multimodal Mixture-of-Experts (DMMoE)
    purpose: Proposed future architecture for adaptive multimodal perception, with a gating network that selects modality or task experts.
    Introduced in Section VI as a forward-looking idea; no implementation, experiments, or falsifiable predictions are provided.
  • Dual knowledge architecture (offline interaction-driven base plus online time-sensitive retrieval)
    purpose: To enable AI agents to integrate embodied experience with up-to-date external information for temporal and causal reasoning.
    Proposed in Section VI as a suggestion, distinct from existing RAG, but not built or tested in the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Nature's Insight: A Novel Framework and Comprehensive Analysis of Agentic Reasoning Through the Lens of Neuroscience." pith.science (2026). https://pith.science/paper/B2YPLKF4

@misc{pith2026250505515,
  author       = {Pith},
  title        = {Pith review of: Nature's Insight: A Novel Framework and Comprehensive Analysis of Agentic Reasoning Through the Lens of Neuroscience},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/B2YPLKF4}},
  note         = {Machine review of arXiv:2505.05515}
}
read the original abstract

Autonomous AI is no longer a hard-to-reach concept, it enables the agents to move beyond executing tasks to independently addressing complex problems, adapting to change while handling the uncertainty of the environment. However, what makes the agents truly autonomous? It is agentic reasoning, that is crucial for foundation models to develop symbolic logic, statistical correlations, or large-scale pattern recognition to process information, draw inferences, and make decisions. However, it remains unclear why and how existing agentic reasoning approaches work, in comparison to biological reasoning, which instead is deeply rooted in neural mechanisms involving hierarchical cognition, multimodal integration, and dynamic interactions. In this work, we propose a novel neuroscience-inspired framework for agentic reasoning. Grounded in three neuroscience-based definitions and supported by mathematical and biological foundations, we propose a unified framework modeling reasoning from perception to action, encompassing four core types, perceptual, dimensional, logical, and interactive, inspired by distinct functional roles observed in the human brain. We apply this framework to systematically classify and analyze existing AI reasoning methods, evaluating their theoretical foundations, computational designs, and practical limitations. We also explore its implications for building more generalizable, cognitively aligned agents in physical and virtual environments. Finally, building on our framework, we outline future directions and propose new neural-inspired reasoning methods, analogous to chain-of-thought prompting. By bridging cognitive neuroscience and AI, this work offers a theoretical foundation and practical roadmap for advancing agentic reasoning in intelligent systems. The associated project can be found at: https://github.com/BioRAILab/Awesome-Neuroscience-Agent-Reasoning .

Figures

Figures reproduced from arXiv: 2505.05515 by the authors.

Figure 1
Figure 1. The proposed neuroscience-inspired framework for agentic reasoning. The left panel illustrates the human brain’s reasoning process, where sensory [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Google Scholar results for research topics related to agentic reasoning. [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. The hybrid nature of reasoning in humans and AI agents. Reasoning is a fusion of prior knowledge and new information, forming a hybrid process. This section provides examples: 1) Human Reasoning, deciding what to wear based on past knowledge and weather forecasts, and 2) Agentic Reasoning, adjusting navigation in response to unexpected obstacles. Multistep Structured Process. Reasoning follows a struc￾tured, multist… view at source ↗
Figures from the paper (14 more)
Figure 4
Figure 4. Figure 4: The overview of the reasoning process and classification of reasoning behavior from a neuro-perspective. This diagram presents a comprehensive [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Taxonomy of Agentic Reasoning Techniques Inspired by Neuroscience. This hierarchical structure organizes reasoning methods in artificial agents based on cognitive mechanisms inspired by neuroscience, including dimensional, perceptual, logical, and interactive reasoning…
Figure 6
Figure 6. Figure 6: Structure of different visual reasoning methods. (a) VLM-based [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: Evolution timeline for LLM lingual reasoning methods. (a): Evolution timeline for CoT-based methods: CoT prompting was first introduced in 2022. [PITH_FULL_IMAGE:figures/full_fig_p015_7.png]
Figure 8
Figure 8. Figure 8: Even with complex step-by-step CoT prompting that reflects how people would actually approach a reasoning problem on a good model, such as [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]
Figure 9
Figure 9. Figure 9: Pipeline for spatial reasoning in object-centric environments. The [PITH_FULL_IMAGE:figures/full_fig_p018_9.png]
Figure 10
Figure 10. Figure 10: Structure of different temporal reasoning methods. (a) and (b) are [PITH_FULL_IMAGE:figures/full_fig_p019_10.png]
Figure 11
Figure 11. Figure 11: The main process of neuro-symbolic learning. Continuous multi [PITH_FULL_IMAGE:figures/full_fig_p020_11.png]
Figure 12
Figure 12. Figure 12: (a), (b) and (c) are the processes in inductive reasoning, deductive reasoning, and abductive reasoning, respectively. We refer to the flowcharts from [PITH_FULL_IMAGE:figures/full_fig_p021_12.png]
Figure 13
Figure 13. Figure 13: Taxonomy of agent-agent and human-agent interaction reasoning across equality and inequality dimensions. This framework categorizes interaction paradigms based on the axis of equality and the nature of interaction, agent-agent versus human-agent. In the top-left quadr…
Figure 14
Figure 14. Figure 14: Representative categories of modern robotic platforms. We showcase four primary types of embodied robotic agents: robot dogs for agile terrain traversal, unmanned aerial vehicles (UAVs) for aerial sensing, humanoid robots designed for human-centric tasks, and robot ma…
Figure 15
Figure 15. Figure 15: An overview of our proposed AI agent system architecture designed to facilitate reasoning through multimodal perception and dynamic knowledge [PITH_FULL_IMAGE:figures/full_fig_p029_15.png]
Figure 16
Figure 16. Figure 16: Framework for continuous spatiotemporal neural differential reasoning in embodied agents. Inspired by the human parietal lobe, this architecture integrates dynamic spatiotemporal event spaces with 3D structural information and time-varying dynamics using Neural ODEs. …
Figure 17
Figure 17. Figure 17: Future AI agents should possess the ability to reason about others [PITH_FULL_IMAGE:figures/full_fig_p031_17.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective

    cs.AI 2025-09 conditional novelty 4.0 of 10

    Agent spatial intelligence is organized into six neuroscience-inspired modules, and the field is reviewed through that lens without any experimental validation.

  2. Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI

    cs.SE 2025-05 conditional novelty 3.0 of 10

    A qualitative taxonomy positions vibe coding and agentic coding as complementary paradigms rather than rivals in AI-assisted software development.

Reference graph

Works this paper leans on

274 extracted references · 2 canonical work pages · cited by 2 Pith papers

  1. [1]

    Routledge, 2009

    Proudfoot, Michael and Lacey, Alan Robert, The Routledge dictionary of philosophy. Routledge, 2009

  2. [2]

    Pearson, 2016

    Russell, Stuart J and Norvig, Peter, Artificial intelligence: a modern approach. Pearson, 2016

  3. [3]

    Artificial intelligence: Structures and strategies for complex problem solving,

    Khorasani, Elham S, “Artificial intelligence: Structures and strategies for complex problem solving,” Scalable Computing: Practice and Experience, vol. 9, no. 3, 2008

  4. [4]

    Computa- tional Intelligence: a logical approach,

    Poole, David and Mackworth, Alan and Goebel, Randy, “Computa- tional Intelligence: a logical approach,” 1998

  5. [5]

    Elsevier, 1998

    Nilsson, Nils J, Artificial intelligence: a new synthesis. Elsevier, 1998

  6. [6]

    Cortico-cortical feedback engages active dendrites in visual cortex,

    Fis ¸ek, Mehmet and Herrmann, Dustin and Egea-Weiss, Alexander and Cloves, Matilda and Bauer, Lisa and Lee, Tai-Ying and Russell, Lloyd E and H ¨ausser, Michael, “Cortico-cortical feedback engages active dendrites in visual cortex,” Nature, vol. 617, no. 7962, pp. 769–776, 2023

  7. [7]

    Relating structure to function: Heschl’s gyrus and acoustic processing,

    Warrier, Catherine and Wong, Patrick and Penhune, Virginia and Zatorre, Robert and Parrish, Todd and Abrams, Daniel and Kraus, Nina, “Relating structure to function: Heschl’s gyrus and acoustic processing,” Journal of Neuroscience, vol. 29, no. 1, pp. 61–69, 2009

  8. [8]

    Feeling our way to machine minds: People’s emotions when perceiving mind in artificial intelli- gence,

    Shank, Daniel B and Graves, Christopher and Gott, Alexander and Gamez, Patrick and Rodriguez, Sophia, “Feeling our way to machine minds: People’s emotions when perceiving mind in artificial intelli- gence,” Computers in Human Behavior , vol. 98, pp. 256–266, 2019

Show all 274 references
  1. [9]

    Prefrontal cortex,

    Fuster, Joaquin M, “Prefrontal cortex,” in Comparative neuroscience and neurobiology. Springer, 2008, pp. 107–109

  2. [10]

    Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects,

    Rao, Rajesh PN and Ballard, Dana H, “Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects,” Nature Neuroscience, vol. 2, no. 1, pp. 79–87, 1999

  3. [11]

    Motor cortex—to act or not to act?

    Ebbesen, Christian Laut and Brecht, Michael, “Motor cortex—to act or not to act?” Nature Reviews Neuroscience, vol. 18, no. 11, pp. 694–705, 2017

  4. [12]

    The working memory costs of a central attentional bottleneck in multitasking,

    Otermans, Pauldy CJ and Parton, Andrew and Szameitat, Andre J, “The working memory costs of a central attentional bottleneck in multitasking,” Psychological Research, vol. 86, no. 6, pp. 1774–1791, 2022

  5. [13]

    A unified attentional bottleneck in the human brain,

    Tombu, Michael N and Asplund, Christopher L and Dux, Paul E and Godwin, Douglass and Martin, Justin W and Marois, Ren ´e, “A unified attentional bottleneck in the human brain,” Proceedings of the National Academy of Sciences , vol. 108, no. 33, pp. 13 426–13 431, 2011

  6. [14]

    Psychology Press, 2014

    Anderson, John R and Lebiere, Christian J, The atomic components of thought. Psychology Press, 2014

  7. [15]

    Chain-of-thought prompting elicits reasoning in large language models,

    Wei, Jason and Wang, Xuezhi and Schuurmans, Dale and Bosma, Maarten and Xia, Fei and Chi, Ed and Le, Quoc V and Zhou, Denny and others, “Chain-of-thought prompting elicits reasoning in large language models,” Advances in Neural Information Processing Systems, vol. 35, pp. 24 8...

  8. [16]

    The distinct modes of vision offered by feedforward and recurrent processing,

    Lamme, Victor AF and Roelfsema, Pieter R, “The distinct modes of vision offered by feedforward and recurrent processing,” Trends in Neurosciences, vol. 23, no. 11, pp. 571–579, 2000

  9. [17]

    Feedforward and feedback interactions between visual cortical areas use different population activity patterns,

    Semedo, Jo ˜ao D and Jasper, Anna I and Zandvakili, Amin and Krishna, Aravind and Aschner, Amir and Machens, Christian K and Kohn, Adam and Yu, Byron M, “Feedforward and feedback interactions between visual cortical areas use different population activity patterns,” Nature Com...

  10. [18]

    Up- dating mental models in predictive reasoning,

    Rodrigo, Mar ´ıa J and Vega, Manuel de and Castaneda, Javier, “Up- dating mental models in predictive reasoning,” European Journal of Cognitive Psychology, vol. 4, no. 2, pp. 141–157, 1992

  11. [19]

    Thinking as Analogy-Making: Toward a Neural Process Account of General Intelligence,

    Holyoak, Keith J., “Thinking as Analogy-Making: Toward a Neural Process Account of General Intelligence,”The Journal of Neuroscience, vol. 45, no. 18, p. e1555242025, 2025

  12. [20]

    A survey of reasoning with foundation models,

    Sun, Jiankai and Zheng, Chuanyang and Xie, Enze and Liu, Zhengying and Chu, Ruihang and Qiu, Jianing and Xu, Jiaqi and Ding, Mingyu and Li, Hongyang and Geng, Mengzhe and others, “A survey of reasoning with foundation models,” arXiv preprint arXiv:2312.11562 , 2023

  13. [21]

    Stop overthinking: A survey on efficient reasoning for large language models,

    Sui, Yang and Chuang, Yu-Neng and Wang, Guanchu and Zhang, Jiamu and Zhang, Tianyi and Yuan, Jiayi and Liu, Hongyi and Wen, Andrew and Chen, Hanjie and Hu, Xia and others, “Stop overthinking: A survey on efficient reasoning for large language models,”arXiv preprint arXiv:2503....

  14. [22]

    From System 1 to System 2: A Survey of Reasoning Large Language Models,

    Li, Zhong-Zhi and Zhang, Duzhen and Zhang, Ming-Liang and Zhang, Jiaxin and Liu, Zengyan and Yao, Yuxuan and Xu, Haotian and Zheng, Junhao and Wang, Pei-Jie and Chen, Xiuyi and others, “From System 1 to System 2: A Survey of Reasoning Large Language Models,” arXiv preprint arX...

  15. [23]

    Multimodal chain-of-thought reasoning: A comprehensive survey,

    Wang, Yaoting and Wu, Shengqiong and Zhang, Yuecheng and Wang, William and Liu, Ziwei and Luo, Jiebo and Fei, Hao, “Multimodal chain-of-thought reasoning: A comprehensive survey,” arXiv preprint arXiv:2503.12605, 2025

  16. [24]

    Why Reasoning Matters? A Survey of Advancements in Multimodal Reasoning (v1),

    Bi, Jing and Liang, Susan and Zhou, Xiaofei and Liu, Pinxin and Guo, Junjia and Tang, Yunlong and Song, Luchuan and Huang, Chao and Sun, Guangyu and He, Jinxi and others, “Why Reasoning Matters? A Survey of Advancements in Multimodal Reasoning (v1),” arXiv preprint arXiv:2504....

  17. [25]

    A case-based reasoner adaptive to different cog- nitive tasks,

    Bichindaritz, Isabelle, “A case-based reasoner adaptive to different cog- nitive tasks,” in International Conference on Case-Based Reasoning . Springer, 1995, pp. 391–400

  18. [26]

    Ontology- oriented case-based reasoning (CBR) approach for trainings adaptive delivery,

    Mansouri, DOUNIA and Hamdi-Cherif, Aboubekeur, “Ontology- oriented case-based reasoning (CBR) approach for trainings adaptive delivery,” in Proceedings of the WSEAS International Conference on Computers, 2011, pp. 328–333

  19. [27]

    Academic Press, 2017

    Krawczyk, Daniel, Reasoning: The neuroscience of how we think . Academic Press, 2017

  20. [28]

    An integrative theory of prefrontal cortex function,

    Miller, Earl K and Cohen, Jonathan D, “An integrative theory of prefrontal cortex function,” Annual Review of Neuroscience , vol. 24, no. 1, pp. 167–202, 2001

  21. [29]

    Executive function: The search for an integrated account,

    Banich, Marie T, “Executive function: The search for an integrated account,” Current directions in psychological science , vol. 18, no. 2, pp. 89–94, 2009

  22. [30]

    Working memory,

    Baddeley, Alan, “Working memory,” Memory, pp. 71–111, 2020

  23. [31]

    The episodic buffer: a new component of working memory?

    ——, “The episodic buffer: a new component of working memory?” Trends in cognitive sciences , vol. 4, no. 11, pp. 417–423, 2000

  24. [32]

    SOAR: An architecture for general intelligence,

    Laird, John E and Newell, Allen and Rosenbloom, Paul S, “SOAR: An architecture for general intelligence,” Artificial Intelligence, vol. 33, no. 1, pp. 1–64, 1987

  25. [33]

    Cambridge University Press, 1993

    Baars, Bernard J, A cognitive theory of consciousness . Cambridge University Press, 1993

  26. [34]

    A neostriatal habit learning system in humans,

    Knowlton, Barbara J and Mangels, Jennifer A and Squire, Larry R, “A neostriatal habit learning system in humans,” Science, vol. 273, no. 5280, pp. 1399–1402, 1996

  27. [35]

    Implicit memory and transformative learning theory: Unconscious cognition,

    Taylor, Edward W, “Implicit memory and transformative learning theory: Unconscious cognition,” in Annual Adult Education Research Conference Proceedings. Oklahoma State University, Occupational and Adult Education, 1997, p. 262

  28. [36]

    The case for implicit category learning,

    Smith, Edward E, “The case for implicit category learning,” Cognitive, Affective, & Behavioral Neuroscience , vol. 8, no. 1, pp. 3–16, 2008

  29. [37]

    Learning strategy differentially impacts memory connections in children and adults,

    Abolghasem, Zahra and Teng, Tiffany H-T and Nexha, Elida and Zhu, Cherrie and Jean, Cindy S and Castrillon, Mariana and Che, Eric and Di Nallo, Eva V and Schlichting, Margaret L, “Learning strategy differentially impacts memory connections in children and adults,” Developmenta...

  30. [38]

    IGI Global, 2017

    Kidd, Terry and Morris Jr, Lonnie R, Handbook of research on instructional systems and educational technology . IGI Global, 2017

  31. [39]

    London, UK: Palgrave Macmillan, 2009, pp

    Sunnev ˚ag, Kjell J., The Impact of New Information . London, UK: Palgrave Macmillan, 2009, pp. 353–374

  32. [40]

    M., Learning Through New Information—A Changing Struc- ture Oriented Approach

    Sandi, A. M., Learning Through New Information—A Changing Struc- ture Oriented Approach. Berlin Heidelberg, Germany: Springer, 1978, pp. 165–166

  33. [41]

    Non- monotonic Reasoning,

    Gerhard Brewka and Ilkka Niemel ¨a and Mirosław Truszczy´nski, “Non- monotonic Reasoning,” in Handbook of Knowledge Representation , ser. Foundations of Artificial Intelligence, Frank van Harmelen and Vladimir Lifschitz and Bruce Porter, Ed. Elsevier, 2008, vol. 3, pp. 239–284

  34. [42]

    The Bayesian brain: the role of uncertainty in neural coding and computation,

    Knill, David C and Pouget, Alexandre, “The Bayesian brain: the role of uncertainty in neural coding and computation,” Trends in Neurosciences, vol. 27, no. 12, pp. 712–719, 2004

  35. [43]

    Bayes in the brain—on Bayesian modelling in neuroscience,

    Colombo, Matteo and Seri `es, Peggy, “Bayes in the brain—on Bayesian modelling in neuroscience,” The British journal for the philosophy of science, 2012

  36. [44]

    Predictive coding,

    Huang, Yanping and Rao, Rajesh PN, “Predictive coding,” Wiley Interdisciplinary Reviews: Cognitive Science , vol. 2, no. 5, pp. 580– 593, 2011

  37. [45]

    With or without you: predic- tive coding and Bayesian inference in the brain,

    Aitchison, Laurence and Lengyel, M ´at´e, “With or without you: predic- tive coding and Bayesian inference in the brain,” Current opinion in neurobiology, vol. 46, pp. 219–227, 2017

  38. [46]

    The free-energy principle: a unified brain theory?

    Friston, Karl, “The free-energy principle: a unified brain theory?” Nature Reviews Neuroscience, vol. 11, no. 2, pp. 127–138, 2010

  39. [47]

    A free energy principle for the brain,

    Friston, Karl and Kilner, James and Harrison, Lee, “A free energy principle for the brain,” Journal of Physiology , vol. 100, no. 1-3, pp. 70–87, 2006. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 34

  40. [48]

    Parietal lobe: from action organization to intention understanding,

    Fogassi, Leonardo and Ferrari, Pier Francesco and Gesierich, Benno and Rozzi, Stefano and Chersi, Fabian and Rizzolatti, Giacomo, “Parietal lobe: from action organization to intention understanding,” Science, vol. 308, no. 5722, pp. 662–667, 2005

  41. [49]

    The hippocampus,

    Knierim, James J, “The hippocampus,” Current Biology, vol. 25, no. 23, pp. R1116–R1121, 2015

  42. [50]

    Structure and function of the cerebral cortex,

    Shipp, Stewart, “Structure and function of the cerebral cortex,” Current Biology, vol. 17, no. 12, pp. R443–R449, 2007

  43. [51]

    Dual-processing accounts of reasoning, judg- ment, and social cognition,

    Evans, Jonathan St BT, “Dual-processing accounts of reasoning, judg- ment, and social cognition,” Annual Review of Psychology , vol. 59, no. 1, pp. 255–278, 2008

  44. [52]

    Building machines that learn and think like people,

    Lake, Brenden M and Ullman, Tomer D and Tenenbaum, Joshua B and Gershman, Samuel J, “Building machines that learn and think like people,” Behavioral and brain sciences , vol. 40, p. e253, 2017

  45. [53]

    Bias in the brain: a diffusion model analysis of prior probability and potential payoff,

    Mulder, Martijn J and Wagenmakers, Eric-Jan and Ratcliff, Roger and Boekel, Wouter and Forstmann, Birte U, “Bias in the brain: a diffusion model analysis of prior probability and potential payoff,” Journal of Neuroscience, vol. 32, no. 7, pp. 2335–2343, 2012

  46. [54]

    Brain networks of perceptual decision-making: an fMRI ALE meta-analysis,

    Keuken, Max C and M ¨uller-Axt, Christa and Langner, Robert and Eickhoff, Simon B and Forstmann, Birte U and Neumann, Jane, “Brain networks of perceptual decision-making: an fMRI ALE meta-analysis,” Frontiers in human neuroscience , vol. 8, p. 445, 2014

  47. [55]

    Springer Science & Business Media, 1998

    Stock, Oliviero, Spatial and temporal reasoning . Springer Science & Business Media, 1998

  48. [56]

    On the dimensionality of Reasoning,

    Kubinger, Klaus D, “On the dimensionality of Reasoning,” Psycho- logical Test and Assessment Modeling , vol. 65, no. 3, pp. 437–447, 2023

  49. [57]

    Geometry and spatial reasoning,

    Clements, Douglas H and Battista, Michael T, “Geometry and spatial reasoning,” Handbook of research on mathematics teaching and learn- ing: A project of the National Council of Teachers of Mathematics , pp. 420–464, 1992

  50. [58]

    Men- tal models and temporal reasoning,

    Schaeken, Walter and Johnson-Laird, PN and d’Ydewalle, Gery, “Men- tal models and temporal reasoning,” Cognition, vol. 60, no. 3, pp. 205– 234, 1996

  51. [59]

    Oxford University Press, 1996

    Allwein, Gerard and Barwise, Jon, Logical reasoning with diagrams . Oxford University Press, 1996

  52. [60]

    Logical reasoning in formal and everyday reasoning tasks,

    Bronkhorst, Hugo and Roorda, Gerrit and Suhre, Cor and Goedhart, Martin, “Logical reasoning in formal and everyday reasoning tasks,” International Journal of Science and Mathematics Education , vol. 18, pp. 1673–1694, 2020

  53. [61]

    GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering,

    Ma, Ziyu and Li, Shutao and Sun, Bin and Cai, Jianfei and Long, Zuxiang and Ma, Fuyan, “GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering,” arXiv preprint arXiv:2402.02503, 2024

  54. [62]

    LISA: Reasoning Segmen- tation via Large Language Model,

    Lai, Xin and Tian, Zhuotao and Chen, Yukang and Li, Yanwei and Yuan, Yuhui and Liu, Shu and Jia, Jiaya, “LISA: Reasoning Segmen- tation via Large Language Model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 9579–9589

  55. [63]

    VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model,

    Shen, Haozhan and Zhang, Zilun and Zhang, Qianqian and Xu, Ruochen and Zhao, Tiancheng, “VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model,” 2025

  56. [64]

    KN-VLM: KNowledge-guided Vision-and- Language Model for visual abductive reasoning,

    Tan, Kuo and Qi, Zhaobo and Zhong, Jianping and Xu, Yuanrong and Zhang, Weigang, “KN-VLM: KNowledge-guided Vision-and- Language Model for visual abductive reasoning,” Multimedia Systems, vol. 31, no. 2, p. 146, 2025

  57. [65]

    Large language models are visual reasoning coordinators,

    Chen, Liangyu and Li, Bo and Shen, Sheng and Yang, Jingkang and Li, Chunyuan and Keutzer, Kurt and Darrell, Trevor and Liu, Ziwei, “Large language models are visual reasoning coordinators,” Advances in Neural Information Processing Systems , vol. 36, pp. 70 115–70 140, 2023

  58. [66]

    Visual chain-of-thought prompting for knowledge-based visual reasoning,

    Chen, Zhenfang and Zhou, Qinhong and Shen, Yikang and Hong, Yining and Sun, Zhiqing and Gutfreund, Dan and Gan, Chuang, “Visual chain-of-thought prompting for knowledge-based visual reasoning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 2, ...

  59. [67]

    Visual CoT: Advancing Multi-Modal Language Models with a Com- prehensive Dataset and Benchmark for Chain-of-Thought Reasoning,

    Shao, Hao and Qian, Shengju and Xiao, Han and Song, Guanglu and Zong, Zhuofan and Wang, Letian and Liu, Yu and Li, Hongsheng, “Visual CoT: Advancing Multi-Modal Language Models with a Com- prehensive Dataset and Benchmark for Chain-of-Thought Reasoning,” Advances in Neural Inf...

  60. [68]

    Enhancing LLM Reasoning via Vision-Augmented Prompting,

    Xiao, Ziyang and Zhang, Dongxiang and Han, Xiongwei and Fu, Xiaojin and Yu, Wing Yin and Zhong, Tao and Wu, Sai and Wang, Yuan and Yin, Jianwei and Chen, Gang, “Enhancing LLM Reasoning via Vision-Augmented Prompting,” Advances in Neural Information Processing Systems, vol. 37,...

  61. [69]

    Visual programming: Compositional visual reasoning without training,

    Gupta, Tanmay and Kembhavi, Aniruddha, “Visual programming: Compositional visual reasoning without training,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 14 953–14 962

  62. [70]

    ExoViP: Step-by-step Verification and Exploration with Exoskele- ton Modules for Compositional Visual Reasoning,

    Wang, Yuxuan and Yuille, Alan and Li, Zhuowan and Zheng, Zilong, “ExoViP: Step-by-step Verification and Exploration with Exoskele- ton Modules for Compositional Visual Reasoning,” arXiv preprint arXiv:2408.02210, 2024

  63. [71]

    ViperGPT: Visual Inference via Python Execution for Reasoning,

    Sur ´ıs, D´ıdac and Menon, Sachit and V ondrick, Carl, “ViperGPT: Visual Inference via Python Execution for Reasoning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 11 888–11 898

  64. [72]

    HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning,

    Ke, Fucai and Cai, Zhixi and Jahangard, Simindokht and Wang, Weiqing and Haghighi, Pari Delir and Rezatofighi, Hamid, “HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning,” in European Conference on Computer Vision . Springer, 2024, pp. 132– 149

  65. [73]

    Vision- R1: Incentivizing Reasoning Capability in Multimodal Large Language Models,

    Huang, Wenxuan and Jia, Bohan and Zhai, Zijie and Cao, Shaosheng and Ye, Zheyu and Zhao, Fei and Hu, Yao and Lin, Shaohui, “Vision- R1: Incentivizing Reasoning Capability in Multimodal Large Language Models,” arXiv preprint arXiv:2503.06749 , 2025

  66. [74]

    Visual-RFT: Visual Reinforcement Fine-Tuning,

    Liu, Ziyu and Sun, Zeyi and Zang, Yuhang and Dong, Xiaoyi and Cao, Yuhang and Duan, Haodong and Lin, Dahua and Wang, Ji- aqi, “Visual-RFT: Visual Reinforcement Fine-Tuning,” arXiv preprint arXiv:2503.01785, 2025

  67. [75]

    MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning,

    Pan, Jiazhen and Liu, Che and Wu, Junde and Liu, Fenglin and Zhu, Jiayuan and Li, Hongwei Bran and Chen, Chen and Ouyang, Cheng and Rueckert, Daniel, “MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning,” arXiv prep...

  68. [76]

    Chain of draft: Thinking faster by writing less,

    Xu, Silei and Xie, Wenhao and Zhao, Lingxiao and He, Pengcheng, “Chain of draft: Thinking faster by writing less,” arXiv preprint arXiv:2502.18600, 2025

  69. [77]

    SmartAgent: Chain-of-User-Thought for Embodied Personalized Agent in Cyber World,

    Zhang, Jiaqi and Gao, Chen and Zhang, Liyuan and Li, Yong and Yin, Hongzhi, “SmartAgent: Chain-of-User-Thought for Embodied Personalized Agent in Cyber World,” arXiv preprint arXiv:2412.07472, 2024

  70. [78]

    Tree of thoughts: Deliberate problem solving with large language models,

    Yao, Shunyu and Yu, Dian and Zhao, Jeffrey and Shafran, Izhak and Griffiths, Tom and Cao, Yuan and Narasimhan, Karthik, “Tree of thoughts: Deliberate problem solving with large language models,” Advances in Neural Information Processing Systems , vol. 36, pp. 11 809–11 822, 2023

  71. [79]

    Graph of thoughts: Solving elaborate problems with large language models,

    Besta, Maciej and Blach, Nils and Kubicek, Ales and Gerstenberger, Robert and Podstawski, Michal and Gianinazzi, Lukas and Gajda, Joanna and Lehmann, Tomasz and Niewiadomski, Hubert and Nyczyk, Piotr and others, “Graph of thoughts: Solving elaborate problems with large languag...

  72. [80]

    Enhancing zero-shot chain-of-thought reasoning in large language models through logic,

    Zhao, Xufeng and Li, Mengdi and Lu, Wenhao and Weber, Cornelius and Lee, Jae Hee and Chu, Kun and Wermter, Stefan, “Enhancing zero-shot chain-of-thought reasoning in large language models through logic,” arXiv preprint arXiv:2309.13339 , 2023

  73. [81]

    Automatic prompt augmentation and selection with chain-of-thought from labeled data,

    Shum, KaShun and Diao, Shizhe and Zhang, Tong, “Automatic prompt augmentation and selection with chain-of-thought from labeled data,” arXiv preprint arXiv:2302.12822 , 2023

  74. [82]

    Active prompting with chain-of-thought for large language models,

    Diao, Shizhe and Wang, Pengcheng and Lin, Yong and Pan, Rui and Liu, Xiang and Zhang, Tong, “Active prompting with chain-of-thought for large language models,” arXiv preprint arXiv:2302.12246 , 2023

  75. [83]

    DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning,

    Guo, Daya and Yang, Dejian and Zhang, Haowei and Song, Junxiao and Zhang, Ruoyu and Xu, Runxin and Zhu, Qihao and Ma, Shirong and Wang, Peiyi and Bi, Xiao and others, “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning,” arXiv preprint arXiv:250...

  76. [84]

    Self-playing adversarial language game enhances LLM reasoning,

    Cheng, Pengyu and Hu, Tianhao and Xu, Han and Zhang, Zhisong and Dai, Yong and Han, Lei and Li, Xiaolong and others, “Self-playing adversarial language game enhances LLM reasoning,” Advances in Neural Information Processing Systems , vol. 37, pp. 126 515–126 543, 2024

  77. [85]

    Making large language models better reasoners with step-aware verifier,

    Li, Yifei and Lin, Zeqi and Zhang, Shizhuo and Fu, Qiang and Chen, Bei and Lou, Jian-Guang and Chen, Weizhu, “Making large language models better reasoners with step-aware verifier,” arXiv preprint arXiv:2206.02336, 2022

  78. [86]

    ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search,

    Zhang, Dan and Zhoubian, Sining and Hu, Ziniu and Yue, Yisong and Dong, Yuxiao and Tang, Jie, “ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search,” Advances in Neural Information Processing Systems, vol. 37, pp. 64 735–64 772, 2024. JOURNAL OF LATEX CLASS FILE...

  79. [87]

    Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations,

    Wang, Peiyi and Li, Lei and Shao, Zhihong and Xu, RX and Dai, Damai and Li, Yifei and Chen, Deli and Wu, Yu and Sui, Zhifang, “Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations,” arXiv preprint arXiv:2312.08935 , 2023

  80. [88]

    Improve mathematical reasoning in language models by automated process supervision,

    Luo, Liangchen and Liu, Yinxiao and Liu, Rosanne and Phatale, Samrat and Lara, Harsh and Li, Yunxuan and Shu, Lei and Zhu, Yun and Meng, Lei and Sun, Jiao and others, “Improve mathematical reasoning in language models by automated process supervision,” arXiv preprint arXiv:240...

  81. [89]

    Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs,

    Lai, Xin and Tian, Zhuotao and Chen, Yukang and Yang, Senqiao and Peng, Xiangru and Jia, Jiaya, “Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs,” arXiv preprint arXiv:2406.18629, 2024

  82. [90]

    Learning Audio Concepts from Counterfactual Natural Language,

    V osoughi, Ali and Bondi, Luca and Wu, Ho-Hsiang and Xu, Chenliang, “Learning Audio Concepts from Counterfactual Natural Language,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2024, pp. 366–370

  83. [91]

    BAT: Learning to Reason about Spatial Sounds with Large Language Models,

    Zheng, Zhisheng and Peng, Puyuan and Ma, Ziyang and Chen, Xie and Choi, Eunsol and Harwath, David, “BAT: Learning to Reason about Spatial Sounds with Large Language Models,” arXiv preprint arXiv:2402.01591, 2024

  84. [92]

    Listen, think, and understand,

    Gong, Yuan and Luo, Hongyin and Liu, Alexander H and Karlinsky, Leonid and Glass, James, “Listen, think, and understand,” arXiv preprint arXiv:2305.10790, 2023

  85. [93]

    Octopi: Object Property Reasoning with Large Tactile- Language Models,

    Yu, Samson and Lin, Kelvin and Xiao, Anxing and Duan, Jiafei and Soh, Harold, “Octopi: Object Property Reasoning with Large Tactile- Language Models,” arXiv preprint arXiv:2405.02794 , 2024

  86. [94]

    TALON: Improving Large Language Model Cognition with Tactility-Vision Fusion,

    Jiang, Xinyi and Wang, Guoming and Li, Huanhuan and Xia, Qinghua and Lu, Rongxing and Tang, Siliang, “TALON: Improving Large Language Model Cognition with Tactility-Vision Fusion,” in IEEE Conference on Industrial Electronics and Applications . IEEE, 2024, pp. 1–6

  87. [95]

    Beyond sight: Finetuning generalist robot policies with heterogeneous sensors via language grounding,

    Jones, Joshua and Mees, Oier and Sferrazza, Carmelo and Stachowicz, Kyle and Abbeel, Pieter and Levine, Sergey, “Beyond sight: Finetuning generalist robot policies with heterogeneous sensors via language grounding,” arXiv preprint arXiv:2501.04693 , 2025

  88. [96]

    Vision-language model-based physical reasoning for robot liquid per- ception,

    Lai, Wenqiang and Zhang, Tianwei and Lam, Tin Lun and Gao, Yuan, “Vision-language model-based physical reasoning for robot liquid per- ception,” in IEEE/RSJ International Conference on Intelligent Robots and Systems. IEEE, 2024, pp. 9652–9659

  89. [97]

    SpatialVLM: En- dowing Vision-Language Models with Spatial Reasoning Capabilities,

    Chen, Boyuan and Xu, Zhuo and Kirmani, Sean and Ichter, Brain and Sadigh, Dorsa and Guibas, Leonidas and Xia, Fei, “SpatialVLM: En- dowing Vision-Language Models with Spatial Reasoning Capabilities,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco...

  90. [98]

    Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs,

    Ranasinghe, Kanchana and Shukla, Satya Narayan and Poursaeed, Omid and Ryoo, Michael S and Lin, Tsung-Yu, “Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition, 2024, pp....

  91. [99]

    SpatialPIN: Enhancing Spatial Reasoning Capa- bilities of Vision-Language Models through Prompting and Interacting 3D Priors,

    Ma, Chenyang and Lu, Kai and Cheng, Ta-Ying and Trigoni, Niki and Markham, Andrew, “SpatialPIN: Enhancing Spatial Reasoning Capa- bilities of Vision-Language Models through Prompting and Interacting 3D Priors,” arXiv preprint arXiv:2403.13438 , 2024

  92. [100]

    StarCraftImage: A Dataset For Prototyping Spatial Reasoning Methods For Multi-Agent Environments,

    Kulinski, Sean and Waytowich, Nicholas R and Hare, James Z and Inouye, David I, “StarCraftImage: A Dataset For Prototyping Spatial Reasoning Methods For Multi-Agent Environments,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition, 2023, pp....

  93. [101]

    GRASP: A Grid-Based Bench- mark for Evaluating Commonsense Spatial Reasoning,

    Tang, Zhisheng and Kejriwal, Mayank, “GRASP: A Grid-Based Bench- mark for Evaluating Commonsense Spatial Reasoning,” arXiv preprint arXiv:2407.01892, 2024

  94. [102]

    Reasoning paths with reference objects elicit quantitative spatial reasoning in large vision-language models,

    Liao, Yuan-Hong and Mahmood, Rafid and Fidler, Sanja and Acuna, David, “Reasoning paths with reference objects elicit quantitative spatial reasoning in large vision-language models,” arXiv preprint arXiv:2409.09788, 2024

  95. [103]

    Structured Spatial Reasoning with Open V ocabulary Object Detectors,

    Nejatishahidin, Negar and V ongala, Madhukar Reddy and Kosecka, Jana, “Structured Spatial Reasoning with Open V ocabulary Object Detectors,” arXiv preprint arXiv:2410.07394 , 2024

  96. [104]

    I Know About

    Meng, Zaiqiao and Zhou, Hao and Chen, Yifang, “I Know About ”Up”! Enhancing Spatial Reasoning in Visual Language Models Through 3D Reconstruction,” arXiv preprint arXiv:2407.14133 , 2024

  97. [105]

    TopV-Nav: Unlocking the Top-View Spatial Reasoning Potential of MLLM for Zero-shot Object Navigation,

    Zhong, Linqing and Gao, Chen and Ding, Zihan and Liao, Yue and Liu, Si, “TopV-Nav: Unlocking the Top-View Spatial Reasoning Potential of MLLM for Zero-shot Object Navigation,” arXiv preprint arXiv:2411.16425, 2024

  98. [106]

    Weakly-Supervised 3D Spatial Reasoning for Text-Based Visual Question Answering,

    Li, Hao and Huang, Jinfa and Jin, Peng and Song, Guoli and Wu, Qi and Chen, Jie, “Weakly-Supervised 3D Spatial Reasoning for Text-Based Visual Question Answering,” IEEE Transactions on Image Processing, vol. 32, pp. 3367–3382, 2023

  99. [107]

    SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning,

    Liu, Yuecheng and Chi, Dafeng and Wu, Shiguang and Zhang, Zhanguang and Hu, Yaochen and Zhang, Lingfeng and Zhang, Yingxue and Wu, Shuang and Cao, Tongtong and Huang, Guowei and oth- ers, “SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Though...

  100. [108]

    End-to-End Navigation with Vision Language Models: Trans- forming Spatial Reasoning into Question-Answering,

    Goetting, Dylan and Singh, Himanshu Gaurav and Loquercio, Anto- nio, “End-to-End Navigation with Vision Language Models: Trans- forming Spatial Reasoning into Question-Answering,” arXiv preprint arXiv:2411.05755, 2024

  101. [109]

    Can brain signals reveal inner alignment with human languages?

    Qiu, Jielin and Han, William and Zhu, Jiacheng and Xu, Mengdi and Weber, Douglas and Li, Bo and Zhao, Ding, “Can brain signals reveal inner alignment with human languages?” in Findings of the Association for Computational Linguistics , 2023, pp. 1789–1804

  102. [110]

    PromptCast: A New Prompt-based Learning Paradigm for Time Series Forecasting,

    Xue, Hao and Salim, Flora D, “PromptCast: A New Prompt-based Learning Paradigm for Time Series Forecasting,” IEEE Transactions on Knowledge and Data Engineering , vol. 36, no. 11, pp. 6851–6864, 2023

  103. [111]

    Large language models can learn temporal reasoning,

    Xiong, Siheng and Payani, Ali and Kompella, Ramana and Fekri, Faramarz, “Large language models can learn temporal reasoning,” arXiv preprint arXiv:2401.06853 , 2024

  104. [112]

    En- hancing temporal sensitivity and reasoning for time-sensitive question answering,

    Yang, Wanqi and Li, Yanda and Fang, Meng and Chen, Ling, “En- hancing temporal sensitivity and reasoning for time-sensitive question answering,” arXiv preprint arXiv:2409.16909 , 2024

  105. [113]

    TempoGPT: Enhancing Temporal Reasoning via Quantizing Embedding,

    Zhang, Haochuan and Yang, Chunhua and Han, Jie and Qin, Liyang and Wang, Xiaoli, “TempoGPT: Enhancing Temporal Reasoning via Quantizing Embedding,” arXiv preprint arXiv:2501.07335 , 2025

  106. [114]

    Text-to-ECG: 12- Lead Electrocardiogram Synthesis conditioned on Clinical Text Re- ports,

    Chung, Hyunseung and Kim, Jiho and Kwon, Joon-Myoung and Jeon, Ki-Hyun and Lee, Min Sung and Choi, Edward, “Text-to-ECG: 12- Lead Electrocardiogram Synthesis conditioned on Clinical Text Re- ports,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signa...

  107. [115]

    Know-evolve: Deep temporal reasoning for dynamic knowledge graphs,

    Trivedi, Rakshit and Dai, Hanjun and Wang, Yichen and Song, Le, “Know-evolve: Deep temporal reasoning for dynamic knowledge graphs,” in International Conference on Machine Learning . PMLR, 2017, pp. 3462–3471

  108. [116]

    Temporal inductive path neural network for temporal knowledge graph reasoning,

    Dong, Hao and Wang, Pengyang and Xiao, Meng and Ning, Zhiyuan and Wang, Pengfei and Zhou, Yuanchun, “Temporal inductive path neural network for temporal knowledge graph reasoning,” Artificial Intelligence, vol. 329, p. 104085, 2024

  109. [117]

    An improving reasoning network for complex question answering over temporal knowledge graphs,

    Jiao, Songlin and Zhu, Zhenfang and Wu, Wenqing and Zuo, Zicheng and Qi, Jiangtao and Wang, Wenling and Zhang, Guangyuan and Liu, Peiyu, “An improving reasoning network for complex question answering over temporal knowledge graphs,” Applied Intelligence , vol. 53, no. 7, pp. 8...

  110. [118]

    Event Graph Guided Compositional Spatial–Temporal Reasoning for Video Question Answering,

    Bai, Ziyi and Wang, Ruiping and Gao, Difei and Chen, Xilin, “Event Graph Guided Compositional Spatial–Temporal Reasoning for Video Question Answering,” IEEE Transactions on Image Processing, vol. 33, pp. 1109–1121, 2024

  111. [119]

    Temporal knowledge graph reasoning with historical contrastive learning,

    Xu, Yi and Ou, Junjie and Xu, Hui and Fu, Luoyi, “Temporal knowledge graph reasoning with historical contrastive learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 4, 2023, pp. 4765–4773

  112. [120]

    THCN: A Hawkes Process Based Temporal Causal Convolutional Network for Extrapolation Reasoning in Temporal Knowledge Graphs,

    Chen, Tingxuan and Long, Jun and Wang, Zidong and Luo, Shuai and Huang, Jincai and Yang, Liu, “THCN: A Hawkes Process Based Temporal Causal Convolutional Network for Extrapolation Reasoning in Temporal Knowledge Graphs,” IEEE Transactions on Knowledge and Data Engineering , 2024

  113. [121]

    Hypothesis search: Inductive reasoning with language models,

    Wang, Ruocheng and Zelikman, Eric and Poesia, Gabriel and Pu, Yewen and Haber, Nick and Goodman, Noah D, “Hypothesis search: Inductive reasoning with language models,” arXiv preprint arXiv:2309.05660, 2023

  114. [122]

    Phenomenal yet puzzling: Testing inductive reasoning capabilities of language models with hypothesis refinement,

    Qiu, Linlu and Jiang, Liwei and Lu, Ximing and Sclar, Melanie and Pyatkin, Valentina and Bhagavatula, Chandra and Wang, Bailin and Kim, Yoon and Choi, Yejin and Dziri, Nouha and others, “Phenomenal yet puzzling: Testing inductive reasoning capabilities of language models with ...

  115. [123]

    Deductive verifica- tion of chain-of-thought reasoning,

    Ling, Zhan and Fang, Yunhao and Li, Xuanlin and Huang, Zhiao and Lee, Mingu and Memisevic, Roland and Su, Hao, “Deductive verifica- tion of chain-of-thought reasoning,” Advances in Neural Information Processing Systems, vol. 36, pp. 36 407–36 433, 2023. JOURNAL OF LATEX CLASS ...

  116. [124]

    Certified deductive reasoning with language models,

    Poesia, Gabriel and Gandhi, Kanishk and Zelikman, Eric and Good- man, Noah D, “Certified deductive reasoning with language models,” arXiv preprint arXiv:2306.04031 , 2023

  117. [125]

    Vi- sual abductive reasoning,

    Liang, Chen and Wang, Wenguan and Zhou, Tianfei and Yang, Yi, “Vi- sual abductive reasoning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 15 565–15 575

  118. [126]

    Multi-modal action chain abductive reasoning,

    Li, Mengze and Wang, Tianbao and Xu, Jiahe and Han, Kairong and Zhang, Shengyu and Zhao, Zhou and Miao, Jiaxu and Zhang, Wenqiao and Pu, Shiliang and Wu, Fei, “Multi-modal action chain abductive reasoning,” in Proceedings of the Annual Meeting of the Association for Computatio...

  119. [127]

    DERA: enhancing large language model completions with dialog-enabled resolving agents,

    Nair, Varun and Schumacher, Elliot and Tso, Geoffrey and Kannan, Anitha, “DERA: enhancing large language model completions with dialog-enabled resolving agents,” arXiv preprint arXiv:2303.17071 , 2023

  120. [128]

    RoCo: Dialectic Multi-Robot Collaboration with Large Language Models,

    Mandi, Zhao and Jain, Shreeya and Song, Shuran, “RoCo: Dialectic Multi-Robot Collaboration with Large Language Models,” in IEEE International Conference on Robotics and Automation . IEEE, 2024, pp. 286–299

  121. [129]

    Chateval: Towards better LLM-based evaluators through multi-agent debate,

    Chan, Chi-Min and Chen, Weize and Su, Yusheng and Yu, Jianxuan and Xue, Wei and Zhang, Shanghang and Fu, Jie and Liu, Zhiyuan, “Chateval: Towards better LLM-based evaluators through multi-agent debate,” arXiv preprint arXiv:2308.07201 , 2023

  122. [130]

    Encouraging divergent thinking in large language models through multi-agent debate,

    Liang, Tian and He, Zhiwei and Jiao, Wenxiang and Wang, Xing and Wang, Yan and Wang, Rui and Yang, Yujiu and Shi, Shuming and Tu, Zhaopeng, “Encouraging divergent thinking in large language models through multi-agent debate,” arXiv preprint arXiv:2305.19118 , 2023

  123. [131]

    A vir- tual conversational agent for teens with autism spectrum disorder: Experimental results and design lessons,

    Ali, Mohammad Rafayet and Razavi, Seyedeh Zahra and Langevin, Raina and Al Mamun, Abdullah and Kane, Benjamin and Rawas- sizadeh, Reza and Schubert, Lenhart K and Hoque, Ehsan, “A vir- tual conversational agent for teens with autism spectrum disorder: Experimental results and ...

  124. [132]

    PEER: A Collaborative Language Model,

    Schick, Timo and Dwivedi-Yu, Jane and Jiang, Zhengbao and Petroni, Fabio and Lewis, Patrick and Izacard, Gautier and You, Qingfei and Nalmpantis, Christoforos and Grave, Edouard and Riedel, Se- bastian, “PEER: A Collaborative Language Model,” arXiv preprint arXiv:2208.11663, 2022

  125. [133]

    SAPIEN: affective virtual agents powered by large language models,

    Hasan, Masum and Ozel, Cengiz and Potter, Sammy and Hoque, Ehsan, “SAPIEN: affective virtual agents powered by large language models,” in International Conference on Affective Computing and Intelligent Interaction Workshops and Demos . IEEE, 2023, pp. 1–3

  126. [134]

    Human-level play in the game of Diplomacy by combining language models with strategic reasoning,

    Meta Fundamental AI Research Diplomacy Team (FAIR)† and Bakhtin, Anton and Brown, Noam and Dinan, Emily and Farina, Gabriele and Flaherty, Colin and Fried, Daniel and Goff, Andrew and Gray, Jonathan and Hu, Hengyuan and others, “Human-level play in the game of Diplomacy by com...

  127. [135]

    Self-consistency improves chain of thought reasoning in language models,

    Wang, Xuezhi and Wei, Jason and Schuurmans, Dale and Le, Quoc and Chi, Ed and Narang, Sharan and Chowdhery, Aakanksha and Zhou, Denny, “Self-consistency improves chain of thought reasoning in language models,” arXiv preprint arXiv:2203.11171 , 2022

  128. [136]

    Large language models are reasoning teachers,

    Ho, Namgyu and Schmid, Laura and Yun, Se-Young, “Large language models are reasoning teachers,” arXiv preprint arXiv:2212.10071 , 2022

  129. [137]

    Abstraction-of-Thought Makes Lan- guage Models Better Reasoners,

    Hong, Ruixin and Zhang, Hongming and Pan, Xiaoman and Yu, Dong and Zhang, Changshui, “Abstraction-of-Thought Makes Lan- guage Models Better Reasoners,” arXiv preprint arXiv:2406.12442 , 2024

  130. [138]

    Chain of code: Reasoning with a language model-augmented code emulator,

    Li, Chengshu and Liang, Jacky and Zeng, Andy and Chen, Xinyun and Hausman, Karol and Sadigh, Dorsa and Levine, Sergey and Fei- Fei, Li and Xia, Fei and Ichter, Brian, “Chain of code: Reasoning with a language model-augmented code emulator,” arXiv preprint arXiv:2312.04474, 2023

  131. [139]

    Interleaved- modal chain-of-thought,

    Gao, Jun and Li, Yongqi and Cao, Ziqiang and Li, Wenjie, “Interleaved- modal chain-of-thought,” arXiv preprint arXiv:2411.19488 , 2024

  132. [140]

    Q&A Prompts: Discovering Rich Visual Clues through Mining Question-Answer Prompts for VQA requiring Diverse World Knowledge,

    Wang, Haibo and Ge, Weifeng, “Q&A Prompts: Discovering Rich Visual Clues through Mining Question-Answer Prompts for VQA requiring Diverse World Knowledge,” in European Conference on Computer Vision. Springer, 2024, pp. 274–292

  133. [141]

    Improving zero-shot visual question answering via large language models with reasoning question prompts,

    Lan, Yunshi and Li, Xiang and Liu, Xin and Li, Yang and Qin, Wei and Qian, Weining, “Improving zero-shot visual question answering via large language models with reasoning question prompts,” in Pro- ceedings of the ACM International Conference on Multimedia , 2023, pp. 4389–4400

  134. [142]

    Visual chain of thought: bridging logical gaps with multimodal infillings,

    Rose, Daniel and Himakunthala, Vaishnavi and Ouyang, Andy and He, Ryan and Mei, Alex and Lu, Yujie and Saxon, Michael and Sonar, Chinmay and Mirza, Diba and Wang, William Yang, “Visual chain of thought: bridging logical gaps with multimodal infillings,” arXiv preprint arXiv:23...

  135. [143]

    End-to-End Chart Summarization via Visual Chain-of-Thought in Vision-Language Models,

    Choi, Raymond and Burns, Frank and Lawrence, Chase, “End-to-End Chart Summarization via Visual Chain-of-Thought in Vision-Language Models,” arXiv preprint arXiv:2502.17589 , 2025

  136. [144]

    LLaV A-o1: Let Vision Language Models Reason Step-by-Step,

    Xu, Guowei and Jin, Peng and Hao, Li and Song, Yibing and Sun, Lichao and Yuan, Li, “LLaV A-o1: Let Vision Language Models Reason Step-by-Step,” arXiv preprint arXiv:2411.10440 , 2024

  137. [145]

    Zero-shot visual reasoning through probabilistic analogical mapping,

    Webb, Taylor and Fu, Shuhao and Bihl, Trevor and Holyoak, Keith J and Lu, Hongjing, “Zero-shot visual reasoning through probabilistic analogical mapping,” Nature Communications, vol. 14, no. 1, p. 5144, 2023

  138. [146]

    VLM-RL: A Unified Vision Language Models and Reinforcement Learning Framework for Safe Autonomous Driving,

    Huang, Zilin and Sheng, Zihao and Qu, Yansong and You, Junwei and Chen, Sikai, “VLM-RL: A Unified Vision Language Models and Reinforcement Learning Framework for Safe Autonomous Driving,” arXiv preprint arXiv:2412.15544 , 2024

  139. [147]

    Dissociating language and thought in human reasoning,

    Coetzee, John P and Johnson, Micah A and Lee, Youngzie and Wu, Allan D and Iacoboni, Marco and Monti, Martin M, “Dissociating language and thought in human reasoning,” Brain Sciences , vol. 13, no. 1, p. 67, 2022

  140. [148]

    PaLM: Scaling Language Modeling with Path- ways,

    Chowdhery, Aakanksha and Narang, Sharan and Devlin, Jacob and Bosma, Maarten and Mishra, Gaurav and Roberts, Adam and Barham, Paul and Chung, Hyung Won and Sutton, Charles and Gehrmann, Sebastian and others, “PaLM: Scaling Language Modeling with Path- ways,” Journal of Machine...

  141. [149]

    Training verifiers to solve math word problems,

    Cobbe, Karl and Kosaraju, Vineet and Bavarian, Mohammad and Chen, Mark and Jun, Heewoo and Kaiser, Lukasz and Plappert, Matthias and Tworek, Jerry and Hilton, Jacob and Nakano, Reiichiro and others, “Training verifiers to solve math word problems,” arXiv preprint arXiv:2110.14...

  142. [150]

    Chain-of-thought reasoning without prompting,

    Wang, Xuezhi and Zhou, Denny, “Chain-of-thought reasoning without prompting,” arXiv preprint arXiv:2402.10200 , 2024

  143. [151]

    To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning,

    Sprague, Zayne and Yin, Fangcong and Rodriguez, Juan Diego and Jiang, Dongwei and Wadhwa, Manya and Singhal, Prasann and Zhao, Xinyu and Ye, Xi and Mahowald, Kyle and Durrett, Greg, “To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning,” arXiv pre...

  144. [152]

    Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents,

    Putta, Pranav and Mills, Edmund and Garg, Naman and Motwani, Sumeet and Finn, Chelsea and Garg, Divyansh and Rafailov, Rafael, “Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents,” arXiv preprint arXiv:2408.07199 , 2024

  145. [153]

    Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning,

    Xie, Tian and Gao, Zitian and Ren, Qingnan and Luo, Haoming and Hong, Yuqian and Dai, Bryan and Zhou, Joey and Qiu, Kai and Wu, Zhirong and Luo, Chong, “Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning,” arXiv preprint arXiv:2502.14768, 2025

  146. [154]

    Reasoning with reinforced functional token tuning,

    Zhang, Kongcheng and Yao, Qi and Lai, Baisheng and Huang, Jiaxing and Fang, Wenkai and Tao, Dacheng and Song, Mingli and Liu, Shunyu, “Reasoning with reinforced functional token tuning,” arXiv preprint arXiv:2502.13389, 2025

  147. [155]

    AST: Audio Spectrogram Transformer,

    Gong, Yuan and Chung, Yu-An and Glass, James, “AST: Audio Spectrogram Transformer,” arXiv preprint arXiv:2104.01778 , 2021

  148. [156]

    LLaMA: Open and Efficient Foundation Language Models,

    Touvron, Hugo and Lavril, Thibaut and Izacard, Gautier and Martinet, Xavier and Lachaux, Marie-Anne and Lacroix, Timoth ´ee and Rozi `ere, Baptiste and Goyal, Naman and Hambro, Eric and Azhar, Faisal and others, “LLaMA: Open and Efficient Foundation Language Models,” arXiv pre...

  149. [157]

    SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models,

    Cheng, An-Chieh and Yin, Hongxu and Fu, Yang and Guo, Qiushan and Yang, Ruihan and Kautz, Jan and Wang, Xiaolong and Liu, Sifei, “SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models,” arXiv preprint arXiv:2406.01584 , 2024

  150. [158]

    Temporal reasoning transfer from text to video,

    Li, Lei and Liu, Yuanxin and Yao, Linli and Zhang, Peiyuan and An, Chenxin and Wang, Lean and Sun, Xu and Kong, Lingpeng and Liu, Qi, “Temporal reasoning transfer from text to video,” arXiv preprint arXiv:2410.06166, 2024

  151. [159]

    Learning spatial models for navi- gation,

    Epstein, Susan L and Aroor, Anoop and Evanusa, Matthew and Sklar, Elizabeth I and Parsons, Simon, “Learning spatial models for navi- gation,” in International Conference on Spatial Information Theory . Springer, 2015, pp. 403–425

  152. [160]

    Artificial Intelligence Geographic Information Systems-AI GIS,

    Ahmed, Zakaria Yehia, “Artificial Intelligence Geographic Information Systems-AI GIS,” International Journal of Advanced Engineering and Business Sciences, vol. 5, no. 1, 2024

  153. [161]

    Deep learn- ing,

    LeCun, Yann and Bengio, Yoshua and Hinton, Geoffrey, “Deep learn- ing,” Nature, vol. 521, no. 7553, pp. 436–444, 2015. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 37

  154. [162]

    The graph neural network model,

    Scarselli, Franco and Gori, Marco and Tsoi, Ah Chung and Hagen- buchner, Markus and Monfardini, Gabriele, “The graph neural network model,” IEEE Transactions on Neural Networks, vol. 20, no. 1, pp. 61– 80, 2008

  155. [163]

    Metric Reasoning in Large Language Models,

    O’Sullivan, Kent and Schneider, Nicole R. and Samet, Hanan, “Metric Reasoning in Large Language Models,” in Proceedings of the ACM International Conference on Advances in Geographic Information Systems, ser. SIGSPATIAL ’24. New York, NY , USA: ACM, 2024, p. 501–504

  156. [164]

    Reframing spa- tial reasoning evaluation in language models: A real-world simulation benchmark for qualitative reasoning,

    Li, Fangjun and Hogg, David C and Cohn, Anthony G, “Reframing spa- tial reasoning evaluation in language models: A real-world simulation benchmark for qualitative reasoning,”arXiv preprint arXiv:2405.15064, 2024

  157. [165]

    Spatial representation and reasoning in RCC-8 with Boolean region terms,

    Wolter, Frank and Zakharyaschev, Michael, “Spatial representation and reasoning in RCC-8 with Boolean region terms,” in Proceedings of the European Conference on Artificial Intelligence . Citeseer, 2000, pp. 244–248

  158. [166]

    Qualitative process theory,

    Forbus, Kenneth D, “Qualitative process theory,” Artificial Intelligence, vol. 24, no. 1-3, pp. 85–168, 1984

  159. [167]

    AZTR: Aerial Video Action Recognition with Auto Zoom and Tem- poral Reasoning,

    Wang, Xijun and Xian, Ruiqi and Guan, Tianrui and de Melo, Celso M and Nogar, Stephen M and Bera, Aniket and Manocha, Dinesh, “AZTR: Aerial Video Action Recognition with Auto Zoom and Tem- poral Reasoning,” in IEEE International Conference on Robotics and Automation. IEEE, 202...

  160. [168]

    ReasonNet: End-to-End Driving with Temporal and Global Reasoning,

    Shao, Hao and Wang, Letian and Chen, Ruobing and Waslander, Steven L and Li, Hongsheng and Liu, Yu, “ReasonNet: End-to-End Driving with Temporal and Global Reasoning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 13 723–13 733

  161. [169]

    Discovering spatio-temporal rationales for video question answering,

    Li, Yicong and Xiao, Junbin and Feng, Chun and Wang, Xiang and Chua, Tat-Seng, “Discovering spatio-temporal rationales for video question answering,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 13 869–13 878

  162. [170]

    3D deformable convolution temporal reasoning network for action recognition,

    Ou, Yangjun and Chen, Zhenzhong, “3D deformable convolution temporal reasoning network for action recognition,” Journal of Visual Communication and Image Representation , vol. 93, p. 103804, 2023

  163. [171]

    TKN: Transformer-based Keypoint Prediction Network For Real-time Video Prediction,

    Li, Haoran and Zhou, Pengyuan and Lin, Yihang and Hao, Yanbin and Xie, Haiyong and Liao, Yong, “TKN: Transformer-based Keypoint Prediction Network For Real-time Video Prediction,” arXiv preprint arXiv:2303.09807, 2023

  164. [172]

    JSTR: Joint Spatio-Temporal Reasoning for Event-based Moving Object Detection,

    Zhou, Hanyu and Shi, Zhiwei and Dong, Hao and Peng, Shihan and Chang, Yi and Yan, Luxin, “JSTR: Joint Spatio-Temporal Reasoning for Event-based Moving Object Detection,” in IEEE International Conference on Robotics and Automation . IEEE, 2024, pp. 10 650– 10 656

  165. [173]

    Finding structure in time,

    Elman, Jeffrey L, “Finding structure in time,” Cognitive science , vol. 14, no. 2, pp. 179–211, 1990

  166. [174]

    Long short-term memory,

    Hochreiter, Sepp and Schmidhuber, J ¨urgen, “Long short-term memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997

  167. [175]

    Learning phrase representations using RNN encoder-decoder for statistical machine translation,

    Cho, Kyunghyun and Van Merri ¨enboer, Bart and Gulcehre, Caglar and Bahdanau, Dzmitry and Bougares, Fethi and Schwenk, Holger and Bengio, Yoshua, “Learning phrase representations using RNN encoder-decoder for statistical machine translation,” arXiv preprint arXiv:1406.1078, 2014

  168. [176]

    Attention is all you need,

    Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N and Kaiser, Łukasz and Polosukhin, Illia, “Attention is all you need,” Advances in Neural Information Processing Systems , vol. 30, 2017

  169. [177]

    Back to the future: Towards explainable temporal reasoning with large language models,

    Yuan, Chenhan and Xie, Qianqian and Huang, Jimin and Ananiadou, Sophia, “Back to the future: Towards explainable temporal reasoning with large language models,” in Proceedings of the ACM Web Confer- ence, 2024, pp. 1963–1974

  170. [178]

    A survey on neural-symbolic learning systems,

    Yu, Dongran and Yang, Bo and Liu, Dayou and Wang, Hui and Pan, Shirui, “A survey on neural-symbolic learning systems,” Neural Networks, vol. 166, pp. 105–126, 2023

  171. [179]

    Gradient-based learning applied to document recognition,

    LeCun, Yann and Bottou, L ´eon and Bengio, Yoshua and Haffner, Patrick, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998

  172. [180]

    Probabilistic logic neural networks for rea- soning,

    Qu, Meng and Tang, Jian, “Probabilistic logic neural networks for rea- soning,” Advances in Neural Information Processing Systems , vol. 32, 2019

  173. [181]

    Efficient proba- bilistic logic reasoning with graph neural networks,

    Zhang, Yuyu and Chen, Xinshi and Yang, Yuan and Ramamurthy, Arun and Li, Bo and Qi, Yuan and Song, Le, “Efficient proba- bilistic logic reasoning with graph neural networks,” arXiv preprint arXiv:2001.11850, 2020

  174. [182]

    Learn to explain efficiently via neural logic inductive learning,

    Yang, Yuan and Song, Le, “Learn to explain efficiently via neural logic inductive learning,” arXiv preprint arXiv:1910.02481 , 2019

  175. [183]

    The neuro-symbolic concept learner: Inter- preting scenes, words, and sentences from natural supervision,

    Mao, Jiayuan and Gan, Chuang and Kohli, Pushmeet and Tenenbaum, Joshua B and Wu, Jiajun, “The neuro-symbolic concept learner: Inter- preting scenes, words, and sentences from natural supervision,” arXiv preprint arXiv:1904.12584, 2019

  176. [184]

    DeepProbLog: Neural Probabilistic Logic Programming,

    Manhaeve, Robin and Dumancic, Sebastijan and Kimmig, Angelika and Demeester, Thomas and De Raedt, Luc, “DeepProbLog: Neural Probabilistic Logic Programming,” Advances in Neural Information Processing Systems, vol. 31, 2018

  177. [185]

    Approx- imate inference for neural probabilistic logic programming,

    Manhaeve, Robin and Marra, Giuseppe and De Raedt, Luc, “Approx- imate inference for neural probabilistic logic programming,” in Pro- ceedings of the International Conference on Principles of Knowledge Representation and Reasoning . IJCAI Organization, 2021, pp. 475– 486

  178. [186]

    Parameter estimation for probabilistic finite-state trans- ducers,

    Eisner, Jason, “Parameter estimation for probabilistic finite-state trans- ducers,” in Proceedings of the Annual Meeting of the Association for Computational Linguistics, 2002, pp. 1–8

  179. [187]

    A probabilistic graphical model based on neural-symbolic reasoning for visual relationship detection,

    Yu, Dongran and Yang, Bo and Wei, Qianhao and Li, Anchen and Pan, Shirui, “A probabilistic graphical model based on neural-symbolic reasoning for visual relationship detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 10...

  180. [188]

    Inductive reasoning in humans and large language models,

    Han, Simon Jerome and Ransom, Keith J and Perfors, Andrew and Kemp, Charles, “Inductive reasoning in humans and large language models,” Cognitive Systems Research , vol. 83, p. 101155, 2024

  181. [189]

    GPT-4 Technical Report,

    Achiam, Josh and Adler, Steven and Agarwal, Sandhini and Ah- mad, Lama and Akkaya, Ilge and Aleman, Florencia Leoni and Almeida, Diogo and Altenschmidt, Janko and Altman, Sam and Anad- kat, Shyamal and others, “GPT-4 Technical Report,” arXiv preprint arXiv:2303.08774, 2023

  182. [190]

    Testing the general deductive reasoning capacity of large language models using ood examples,

    Saparov, Abulhair and Pang, Richard Yuanzhe and Padmakumar, Vishakh and Joshi, Nitish and Kazemi, Mehran and Kim, Najoung and He, He, “Testing the general deductive reasoning capacity of large language models using ood examples,” Advances in Neural Information Processing Syste...

  183. [191]

    CaPo: Cooperative Plan Optimization for Efficient Embodied Multi-Agent Cooperation,

    Liu, Jie and Zhou, Pan and Du, Yingjun and Tan, Ah-Hwee and Snoek, Cees GM and Sonke, Jan-Jakob and Gavves, Efstratios, “CaPo: Cooperative Plan Optimization for Efficient Embodied Multi-Agent Cooperation,” arXiv preprint arXiv:2411.04679 , 2024

  184. [192]

    Building cooperative embodied agents modularly with large language models,

    Zhang, Hongxin and Du, Weihua and Shan, Jiaming and Zhou, Qin- hong and Du, Yilun and Tenenbaum, Joshua B and Shu, Tianmin and Gan, Chuang, “Building cooperative embodied agents modularly with large language models,” arXiv preprint arXiv:2307.02485 , 2023

  185. [193]

    Language grounded multi-agent reinforcement learning with human-interpretable communication,

    Li, Huao and Nourkhiz Mahjoub, Hossein and Chalaki, Behdad and Tadiparthi, Vaishnav and Lee, Kwonjoon and Moradi Pari, Ehsan and Lewis, Charles and Sycara, Katia, “Language grounded multi-agent reinforcement learning with human-interpretable communication,” Ad- vances in Neura...

  186. [194]

    Simon and Schuster, 1988

    Minsky, Marvin, Society of mind . Simon and Schuster, 1988

  187. [195]

    VQA: Visual Question Answering,

    Antol, Stanislaw and Agrawal, Aishwarya and Lu, Jiasen and Mitchell, Margaret and Batra, Dhruv and Zitnick, C Lawrence and Parikh, Devi, “VQA: Visual Question Answering,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2015, pp. 2425–2433

  188. [196]

    Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering,

    Goyal, Yash and Khot, Tejas and Summers-Stay, Douglas and Batra, Dhruv and Parikh, Devi, “Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, ...

  189. [197]

    Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning,

    Li, Zhuowan and Wang, Xingrui and Stengel-Eskin, Elias and Ko- rtylewski, Adam and Ma, Wufei and Van Durme, Benjamin and Yuille, Alan L, “Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning,” in Proceedings of the IEEE/CVF Conference on Computer ...

  190. [198]

    GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answer- ing,

    Hudson, Drew A and Manning, Christopher D, “GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answer- ing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 6700–6709

  191. [199]

    A corpus for reasoning about natural language grounded in photographs,

    Suhr, Alane and Zhou, Stephanie and Zhang, Ally and Zhang, Iris and Bai, Huajun and Artzi, Yoav, “A corpus for reasoning about natural language grounded in photographs,” arXiv preprint arXiv:1811.00491, 2018

  192. [200]

    OK-VQA: A Visual Question Answering Bench- mark Requiring External Knowledge,

    Marino, Kenneth and Rastegari, Mohammad and Farhadi, Ali and Mottaghi, Roozbeh, “OK-VQA: A Visual Question Answering Bench- mark Requiring External Knowledge,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 3195–3204. JOURNAL O...

  193. [201]

    A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge,

    Schwenk, Dustin and Khandelwal, Apoorv and Clark, Christopher and Marino, Kenneth and Mottaghi, Roozbeh, “A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge,” in European Conference on Computer Vision . Springer, 2022, pp. 146–162

  194. [202]

    MR-Ben: A Meta- Reasoning Benchmark for Evaluating System-2 Thinking in LLMs,

    Zhongshen Zeng and Yinhong Liu and Yingjia Wan and Jingyao Li and Pengguang Chen and Jianbo Dai and Yuxuan Yao and Rongwu Xu and Zehan Qi and Wanru Zhao and Linling Shen and Jianqiao Lu and Haochen Tan and Yukang Chen and Hao Zhang and Zhan Shi and Bailin Wang and Zhijiang Guo...

  195. [203]

    RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style,

    Yantao Liu and Zijun Yao and Rui Min and Yixin Cao and Lei Hou and Juanzi Li, “RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style,” inInternational Conference on Learning Representations , 2025

  196. [204]

    LR 2 Bench: Evaluating Long-chain Reflective Reasoning Capabilities of Large Language Models via Constraint Satisfaction Problems,

    Chen, Jianghao and Wei, Zhenlin and Ren, Zhenjiang and Li, Ziyong and Zhang, Jiajun, “LR 2 Bench: Evaluating Long-chain Reflective Reasoning Capabilities of Large Language Models via Constraint Satisfaction Problems,” arXiv preprint arXiv:2502.17848 , 2025

  197. [205]

    Big- Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models,

    Albalak, Alon and Phung, Duy and Lile, Nathan and Rafailov, Rafael and Gandhi, Kanishk and Castricato, Louis and Singh, Anikait and Blagden, Chase and Xiang, Violet and Mahan, Dakota and others, “Big- Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in...

  198. [206]

    LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion,

    Ling, Zhan and Liu, Kang and Yan, Kai and Yang, Yifan and Lin, Weijian and Fan, Ting-Han and Shen, Lingfeng and Du, Zhengyin and Chen, Jiecao, “LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion,” arXiv preprint arXiv:2501.15089, 2025

  199. [207]

    Big-bench extra hard,

    Kazemi, Mehran and Fatemi, Bahare and Bansal, Hritik and Palowitch, John and Anastasiou, Chrysovalantis and Mehta, Sanket Vaibhav and Jain, Lalit K and Aglietti, Virginia and Jindal, Disha and Chen, Peter and others, “Big-bench extra hard,” arXiv preprint arXiv:2502.19187 , 2025

  200. [208]

    ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition,

    Liu, Yujie and Yang, Zonglin and Xie, Tong and Ni, Jinjie and Gao, Ben and Li, Yuqiang and Tang, Shixiang and Ouyang, Wanli and Cambria, Erik and Zhou, Dongzhan, “ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition,” arXiv preprint...

  201. [209]

    MastermindEval: A Simple But Scalable Reasoning Benchmark,

    Golde, Jonas and Haller, Patrick and Barth, Fabio and Akbik, Alan, “MastermindEval: A Simple But Scalable Reasoning Benchmark,” arXiv preprint arXiv:2503.05891 , 2025

  202. [210]

    Z1: Efficient Test-time Scaling with Code,

    Yu, Zhaojian and Wu, Yinghao and Zhao, Yilun and Cohan, Arman and Zhang, Xiao-Ping, “Z1: Efficient Test-time Scaling with Code,” arXiv preprint arXiv:2504.00810 , 2025

  203. [211]

    AudioCaps: Generating Captions for Audios in The Wild,

    Kim, Chris Dongjoo and Kim, Byeongchang and Lee, Hyunmin and Kim, Gunhee, “AudioCaps: Generating Captions for Audios in The Wild,” inProceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2019,...

  204. [212]

    Clotho: An Audio Captioning Dataset,

    Drossos, Konstantinos and Lipping, Samuel and Virtanen, Tuomas, “Clotho: An Audio Captioning Dataset,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2020, pp. 736–740

  205. [213]

    Transferable tactile transformers for representation learning across diverse sensors and tasks,

    Zhao, Jialiang and Ma, Yuxiang and Wang, Lirui and Adelson, Edward H, “Transferable tactile transformers for representation learning across diverse sensors and tasks,” arXiv preprint arXiv:2406.13640 , 2024

  206. [214]

    Touch100k: A large-scale touch- language-vision dataset for touch-centric multimodal representation,

    Cheng, Ning and Guan, Changhao and Gao, Jing and Wang, Weihao and Li, You and Meng, Fandong and Zhou, Jie and Fang, Bin and Xu, Jinan and Han, Wenjuan, “Touch100k: A large-scale touch- language-vision dataset for touch-centric multimodal representation,” arXiv preprint arXiv:2...

  207. [215]

    Any- Touch: Learning Unified Static-Dynamic Representation across Mul- tiple Visuo-tactile Sensors,

    Feng, Ruoxuan and Hu, Jiangyu and Xia, Wenke and Gao, Tianci and Shen, Ao and Sun, Yuhao and Fang, Bin and Hu, Di, “Any- Touch: Learning Unified Static-Dynamic Representation across Mul- tiple Visuo-tactile Sensors,” arXiv preprint arXiv:2502.12191 , 2025

  208. [216]

    RA VEN: A Dataset for Relational and Analogical Visual rEasoNing,

    Zhang, Chi and Gao, Feng and Jia, Baoxiong and Lu, Jiajun and Zhu, Song-Chun, “RA VEN: A Dataset for Relational and Analogical Visual rEasoNing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019

  209. [217]

    SPARQA: A Spatial Reasoning Question Answering Dataset for Visual Scene Understanding,

    Perez, Ethan and Kembhavi, Aniruddha and Zitnick, C Lawrence and Farhadi, Ali and Hajishirzi, Hannaneh, “SPARQA: A Spatial Reasoning Question Answering Dataset for Visual Scene Understanding,” in Findings of the Association for Computational Linguistics , 2021

  210. [218]

    GRiT: General Robust Image Task Benchmark for Spatial Graph Reasoning,

    Yang, Xiaojian and Li, Yuncheng and Wang, Xin and Darrell, Trevor, “GRiT: General Robust Image Task Benchmark for Spatial Graph Reasoning,” Advances in Neural Information Processing Systems, 2022

  211. [219]

    You need to pay attention: Fine-grained visual question answering,

    Kembhavi, Aniruddha and Salvato, Tejas and Kolve, Eric and et al., “You need to pay attention: Fine-grained visual question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2017

  212. [220]

    CoDraw: Collaborative Drawing as a Testbed for Grounded Goal-driven Communication,

    Kim, Jae Sung and et al., “CoDraw: Collaborative Drawing as a Testbed for Grounded Goal-driven Communication,” in Proceedings of the Annual Meeting of the Association for Computational Linguistics , 2019

  213. [221]

    Touch- down: Natural language navigation and spatial reasoning in visual street environments,

    Chen, Howard and Suhr, Alane and Misra, Dipendra and et al., “Touch- down: Natural language navigation and spatial reasoning in visual street environments,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019

  214. [222]

    Vision- and-language navigation: Interpreting visually-grounded navigation in- structions in real environments,

    Anderson, Peter and Wu, Qi and Teney, Damien and et al., “Vision- and-language navigation: Interpreting visually-grounded navigation in- structions in real environments,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2018

  215. [223]

    SpatialSense: An Adversarially Crowdsourced Benchmark for Spatial Relation Recognition,

    Yang, Yi-Lin and Zellers, Rowan and Farhadi, Ali and Choi, Yejin, “SpatialSense: An Adversarially Crowdsourced Benchmark for Spatial Relation Recognition,” in Findings of the Association for Computa- tional Linguistics, 2019

  216. [224]

    A dataset for answering time-sensitive questions,

    Chen, Wenhu and Wang, Xinyi and Wang, William Yang, “A dataset for answering time-sensitive questions,” arXiv preprint arXiv:2108.06314, 2021

  217. [225]

    Time-aware language models as temporal knowledge bases,

    Dhingra, Bhuwan and Cole, Jeremy R and Eisenschlos, Julian Martin and Gillick, Daniel and Eisenstein, Jacob and Cohen, William W, “Time-aware language models as temporal knowledge bases,” Trans- actions of the Association for Computational Linguistics , vol. 10, pp. 257–273, 2022

  218. [226]

    StreamingQA: A Benchmark for Adaptation to New Knowledge over Time in Question Answering Models,

    Liska, Adam and Kocisky, Tomas and Gribovskaya, Elena and Terzi, Tayfun and Sezener, Eren and Agrawal, Devang and D’Autume, Cyprien De Masson and Scholtes, Tim and Zaheer, Manzil and Young, Susannah and others, “StreamingQA: A Benchmark for Adaptation to New Knowledge over Tim...

  219. [227]

    Towards bench- marking and improving the temporal reasoning capability of large language models,

    Tan, Qingyu and Ng, Hwee Tou and Bing, Lidong, “Towards bench- marking and improving the temporal reasoning capability of large language models,” arXiv preprint arXiv:2306.08952 , 2023

  220. [228]

    MenatQA: A New Dataset for Testing the Temporal Comprehension and Reasoning Abilities of Large Language Models,

    Wei, Yifan and Su, Yisong and Ma, Huanhuan and Yu, Xiaoyan and Lei, Fangyu and Zhang, Yuanzhe and Zhao, Jun and Liu, Kang, “MenatQA: A New Dataset for Testing the Temporal Comprehension and Reasoning Abilities of Large Language Models,” arXiv preprint arXiv:2310.05157, 2023

  221. [229]

    TRAM: Benchmarking Temporal Rea- soning for Large Language Models,

    Wang, Yuqing and Zhao, Yun, “TRAM: Benchmarking Temporal Rea- soning for Large Language Models,” arXiv preprint arXiv:2310.00835, 2023

  222. [230]

    ReClor: A Reading Comprehension Dataset Requiring Logical Rea- soning,

    Yu, Weihao and Jiang, Zihang and Dong, Yanfei and Feng, Jiashi, “ReClor: A Reading Comprehension Dataset Requiring Logical Rea- soning,” arXiv preprint arXiv:2002.04326 , 2020

  223. [231]

    Diagnosing the first-order logical reasoning ability through LogicNLI,

    Tian, Jidong and Li, Yitian and Chen, Wenqing and Xiao, Liqiang and He, Hao and Jin, Yaohui, “Diagnosing the first-order logical reasoning ability through LogicNLI,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing , 2021, pp. 3738–3747

  224. [232]

    FOLIO: Natural Language Reasoning with First-Order Logic,

    Han, Simeng and Schoelkopf, Hailey and Zhao, Yilun and Qi, Zhent- ing and Riddell, Martin and Zhou, Wenfei and Coady, James and Peng, David and Qiao, Yujie and Benson, Luke and others, “FOLIO: Natural Language Reasoning with First-Order Logic,” arXiv preprint arXiv:2209.00840, 2022

  225. [233]

    From lsat: The progress and challenges of complex reasoning,

    Wang, Siyuan and Liu, Zhongkun and Zhong, Wanjun and Zhou, Ming and Wei, Zhongyu and Chen, Zhumin and Duan, Nan, “From lsat: The progress and challenges of complex reasoning,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 2201–2216, 2022

  226. [234]

    LogiQA 2.0—An Improved Dataset for Logical Reasoning in Natural Language Under- standing,

    Liu, Hanmeng and Liu, Jian and Cui, Leyang and Teng, Zhiyang and Duan, Nan and Zhou, Ming and Zhang, Yue, “LogiQA 2.0—An Improved Dataset for Logical Reasoning in Natural Language Under- standing,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, pp. 2...

  227. [235]

    Log- icBench: A Benchmark for Evaluation of Logical Reasoning,

    Parmar, Mihir and Varshney, Neeraj and Patel, Nisarg and Mashetty, Santosh and Luo, Man and Mitra, Arindam and Baral, Chitta, “Log- icBench: A Benchmark for Evaluation of Logical Reasoning,” 2023

  228. [236]

    LINGOLY: A Benchmark of Olympiad-Level Linguis- tic Reasoning Puzzles in Low-Resource and Extinct Languages,

    Bean, Andrew M and Hellsten, Simi and Mayne, Harry and Magomere, Jabez and Chi, Ethan A and Chi, Ryan and Hale, Scott A and Kirk, JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 39 Hannah Rose, “LINGOLY: A Benchmark of Olympiad-Level Linguis- tic Reasoning Puzzles in...

  229. [237]

    Roses are red, violets are blue... but should VQA expect them to?

    Kervadec, Corentin and Antipov, Grigory and Baccouche, Moez and Wolf, Christian, “Roses are red, violets are blue... but should VQA expect them to?” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 2776–2785

  230. [238]

    CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning,

    Johnson, Justin and Hariharan, Bharath and Van Der Maaten, Laurens and Fei-Fei, Li and Lawrence Zitnick, C and Girshick, Ross, “CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning,” in Proceedings of the IEEE/CVF Conference on Computer Vision...

  231. [239]

    Navigation through unknown and dynamic open spaces using topological notions,

    Miguel-Tom ´e, Sergio, “Navigation through unknown and dynamic open spaces using topological notions,” Connection Science, vol. 30, no. 2, pp. 160–185, 2018

  232. [240]

    Spatial representation and reasoning for human-robot collaboration,

    Kennedy, William G and Bugajska, Magdalena D and Marge, Matthew and Adams, William and Fransen, Benjamin R and Perzanowski, Den- nis and Schultz, Alan C and Trafton, J Gregory, “Spatial representation and reasoning for human-robot collaboration,” in Proceedings of the AAAI Con...

  233. [241]

    Integrated commonsense reasoning and deep learning for transparent decision making in robotics,

    Mota, Tiago and Sridharan, Mohan and Leonardis, Ale ˇs, “Integrated commonsense reasoning and deep learning for transparent decision making in robotics,” SN Computer Science, vol. 2, no. 4, p. 242, 2021

  234. [242]

    Trust-aware decision making for human- robot collaboration: Model learning and planning,

    Chen, Min and Nikolaidis, Stefanos and Soh, Harold and Hsu, David and Srinivasa, Siddhartha, “Trust-aware decision making for human- robot collaboration: Model learning and planning,” ACM Transactions on Human-Robot Interaction , vol. 9, no. 2, pp. 1–23, 2020

  235. [243]

    Vision AI-based human-robot collaborative assembly driven by autonomous robots,

    Liu, Sichao and Zhang, Jianjing and Wang, Lihui and Gao, Robert X, “Vision AI-based human-robot collaborative assembly driven by autonomous robots,” CIRP annals, vol. 73, no. 1, pp. 13–16, 2024

  236. [244]

    Embodied artificial intelligence: Trends and challenges,

    Pfeifer, Rolf and Iida, Fumiya, “Embodied artificial intelligence: Trends and challenges,” Lecture notes in computer science , pp. 1–26, 2004

  237. [245]

    Visuomotor navigation for embodied robots with spatial memory and semantic reasoning cognition,

    Liu, Qiming and Wang, Guangzhan and Liu, Zhe and Wang, Hesheng, “Visuomotor navigation for embodied robots with spatial memory and semantic reasoning cognition,” IEEE Transactions on Neural Networks and Learning Systems , 2024

  238. [246]

    Hazard challenge: Embodied deci- sion making in dynamically changing environments,

    Zhou, Qinhong and Chen, Sunli and Wang, Yisong and Xu, Haozhe and Du, Weihua and Zhang, Hongxin and Du, Yilun and Tenenbaum, Joshua B and Gan, Chuang, “Hazard challenge: Embodied deci- sion making in dynamically changing environments,” arXiv preprint arXiv:2401.12975, 2024

  239. [247]

    Sensory gain control (amplification) as a mechanism of selective attention: electrophysiological and neuroimaging evidence,

    Hillyard, Steven A and V ogel, Edward K and Luck, Steven J, “Sensory gain control (amplification) as a mechanism of selective attention: electrophysiological and neuroimaging evidence,” Philosophical Trans- actions of the Royal Society of London. Series B: Biological Sciences ...

  240. [248]

    Hearing in complex environments: auditory gain control, attention, and hearing loss,

    Auerbach, Benjamin D and Gritton, Howard J, “Hearing in complex environments: auditory gain control, attention, and hearing loss,” Fron- tiers in neuroscience , vol. 16, p. 799787, 2022

  241. [249]

    Spiking neural net- works,

    Ghosh-Dastidar, Samanwoy and Adeli, Hojjat, “Spiking neural net- works,” International Journal of Neural Systems , vol. 19, no. 04, pp. 295–308, 2009

  242. [250]

    Retrieval-augmented generation for large language models: A survey,

    Gao, Yunfan and Xiong, Yun and Gao, Xinyu and Jia, Kangxiang and Pan, Jinliu and Bi, Yuxi and Dai, Yi and Sun, Jiawei and Wang, Haofen and Wang, Haofen, “Retrieval-augmented generation for large language models: A survey,” arXiv preprint arXiv:2312.10997 , vol. 2, 2023

  243. [251]

    Qwen2.5 Technical Report,

    Yang, An and Yang, Baosong and Zhang, Beichen and Hui, Binyuan and Zheng, Bo and Yu, Bowen and Li, Chengyuan and Liu, Dayiheng and Huang, Fei and Wei, Haoran and others, “Qwen2.5 Technical Report,” arXiv preprint arXiv:2412.15115 , 2024

  244. [252]

    Qwen-VL: A Versatile Vision-Language Model for Un- derstanding, Localization, Text Reading, and Beyond,

    Jinze Bai and Shuai Bai and Shusheng Yang and Shijie Wang and Sinan Tan and Peng Wang and Junyang Lin and Chang Zhou and Jingren Zhou, “Qwen-VL: A Versatile Vision-Language Model for Un- derstanding, Localization, Text Reading, and Beyond,” arXiv preprint arXiv:2308.12966, 2023

  245. [253]

    GPT- 4o System Card,

    Hurst, Aaron and Lerer, Adam and Goucher, Adam P and Perelman, Adam and Ramesh, Aditya and Clark, Aidan and Ostrow, AJ and Welihinda, Akila and Hayes, Alan and Radford, Alec and others, “GPT- 4o System Card,” arXiv preprint arXiv:2410.21276 , 2024

  246. [254]

    The parahippocampal place area: recognition, naviga- tion, or encoding?

    Epstein, Russell and Harris, Alison and Stanley, Damian and Kan- wisher, Nancy, “The parahippocampal place area: recognition, naviga- tion, or encoding?” Neuron, vol. 23, no. 1, pp. 115–125, 1999

  247. [255]

    A cortical representation of the local visual environment,

    Epstein, Russell and Kanwisher, Nancy, “A cortical representation of the local visual environment,” Nature, vol. 392, no. 6676, pp. 598–601, 1998

  248. [256]

    Parahippocampal and retrosplenial contributions to human spatial navigation,

    Epstein, Russell A, “Parahippocampal and retrosplenial contributions to human spatial navigation,” Trends in cognitive sciences , vol. 12, no. 10, pp. 388–396, 2008

  249. [257]

    Human hip- pocampal and entorhinal neurons encode the temporal structure of experience,

    P. Tacikowski, G. Kalender, D. Ciliberti, and I. Fried, “Human hip- pocampal and entorhinal neurons encode the temporal structure of experience,” Nature, vol. 635, no. 8037, pp. 160–167, 2024

  250. [258]

    The cognitive map in humans: spatial navigation and beyond,

    R. A. Epstein, E. Z. Patai, J. B. Julian, and H. J. Spiers, “The cognitive map in humans: spatial navigation and beyond,” Nature neuroscience, vol. 20, no. 11, pp. 1504–1513, 2017

  251. [259]

    4D Gaussian Splatting for Real-Time Dynamic Scene Rendering,

    Wu, Guanjun and Yi, Taoran and Fang, Jiemin and Xie, Lingxi and Zhang, Xiaopeng and Wei, Wei and Liu, Wenyu and Tian, Qi and Wang, Xinggang, “4D Gaussian Splatting for Real-Time Dynamic Scene Rendering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern ...

  252. [260]

    Align Your Gaussians: Text-to-4D with Dynamic 3D Gaussians and Composed Diffusion Models,

    Ling, Huan and Kim, Seung Wook and Torralba, Antonio and Fidler, Sanja and Kreis, Karsten, “Align Your Gaussians: Text-to-4D with Dynamic 3D Gaussians and Composed Diffusion Models,” in Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, ...

  253. [261]

    Neural ordinary differential equations,

    Chen, Ricky TQ and Rubanova, Yulia and Bettencourt, Jesse and Duvenaud, David K, “Neural ordinary differential equations,”Advances in Neural Information Processing Systems , vol. 31, 2018

  254. [262]

    Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training,

    Yuan, Siyu and Chen, Zehui and Xi, Zhiheng and Ye, Junjie and Du, Zhengyin and Chen, Jiecao, “Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training,” arXiv preprint arXiv:2501.11425, 2025

  255. [263]

    Vision-language models can self- improve reasoning via reflection,

    Cheng, Kanzhi and Li, Yantao and Xu, Fangzhi and Zhang, Jianbing and Zhou, Hao and Liu, Yang, “Vision-language models can self- improve reasoning via reflection,” arXiv preprint arXiv:2411.00855 , 2024

  256. [264]

    Meta-Reflection: A Feedback-Free Reflection Learning Framework,

    Wang, Yaoke and Zhu, Yun and Bao, Xintong and Zhang, Wenqiao and Dai, Suyang and Chen, Kehan and Li, Wenqiang and Huang, Gang and Tang, Siliang and Zhuang, Yueting, “Meta-Reflection: A Feedback-Free Reflection Learning Framework,” arXiv preprint arXiv:2412.13781 , 2024

  257. [265]

    Cascaded attention: Adaptive and gated graph attention network for multiagent reinforcement learning,

    Qi, Shuhan and Huang, Xinhao and Peng, Peixi and Huang, Xuzhong and Zhang, Jiajia and Wang, Xuan, “Cascaded attention: Adaptive and gated graph attention network for multiagent reinforcement learning,” IEEE Transactions on Neural Networks and Learning Systems, vol. 35, no. 3, ...

  258. [266]

    Bidirectional cascaded multimodal attention for multiple choice visual question answering,

    Upadhyay, Sushmita and Tripathy, Sanjaya Shankar, “Bidirectional cascaded multimodal attention for multiple choice visual question answering,” Machine Vision and Applications , vol. 36, no. 2, p. 41, 2025

  259. [267]

    Multi-Stage Production Decisions Based on Monte Carlo and Markov Decision Algorithms,

    Zhao, Zhihao and Zhao, Jiahe and Wang, Jiaqi and Xu, Bingrui and Gao, Heyu and Liu, Ji, “Multi-Stage Production Decisions Based on Monte Carlo and Markov Decision Algorithms,” in International Conference on Data Analytics, Computing and Artificial Intelligence . IEEE, 2024, pp...

  260. [268]

    The Buffer Mechanism for Multi-Step Information Reasoning in Language Mod- els,

    Wang, Zhiwei and Wang, Yunji and Zhang, Zhongwang and Zhou, Zhangchen and Jin, Hui and Hu, Tianyang and Sun, Jiacheng and Li, Zhenguo and Zhang, Yaoyu and Xu, Zhi-Qin John, “The Buffer Mechanism for Multi-Step Information Reasoning in Language Mod- els,” arXiv preprint arXiv:2...

  261. [269]

    Central attention is serial, but midlevel and peripheral attention are parallel—A hypothe- sis,

    Tamber-Rosenau, Benjamin J and Marois, Ren ´e, “Central attention is serial, but midlevel and peripheral attention are parallel—A hypothe- sis,” Attention, Perception, & Psychophysics , vol. 78, pp. 1874–1888, 2016

  262. [270]

    Brain mechanisms of serial and parallel processing during dual-task performance,

    Sigman, Mariano and Dehaene, Stanislas, “Brain mechanisms of serial and parallel processing during dual-task performance,” Journal of Neuroscience, vol. 28, no. 30, pp. 7585–7598, 2008

  263. [271]

    SelFee: Iterative Self- Revising LLM Empowered by Self-Feedback Generation,

    Ye, Seonghyeon and Jo, Yongrae and Kim, Doyoung and Kim, Sung- dong and Hwang, Hyeonbin and Seo, Minjoon, “SelFee: Iterative Self- Revising LLM Empowered by Self-Feedback Generation,” Blog post, 2023

  264. [272]

    Visual cognition in multimodal large language models,

    Schulze Buschoff, Luca M and Akata, Elif and Bethge, Matthias and Schulz, Eric, “Visual cognition in multimodal large language models,” Nature Machine Intelligence , pp. 1–11, 2025

  265. [273]

    To- kenCarve: Information-Preserving Visual Token Compression in Mul- timodal Large Language Models,

    Tan, Xudong and Ye, Peng and Tu, Chongjun and Cao, Jianjian and Yang, Yaoxin and Zhang, Lin and Zhou, Dongzhan and Chen, Tao, “To- kenCarve: Information-Preserving Visual Token Compression in Mul- timodal Large Language Models,” arXiv preprint arXiv:2503.10501 , 2025

  266. [274]

    Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model,

    Liu, Ting and Shi, Liangtao and Hong, Richang and Hu, Yue and Yin, Quanjun and Zhang, Linfeng, “Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model,” arXiv preprint arXiv:2411.10803, 2024

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.