Pith. sign in

REVIEW 6 cited by

AlphaMaze: Enhancing Large Language Models' Spatial Intelligence via GRPO

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.14669 v3 pith:DHFJL7GB submitted 2025-02-20 cs.CL

classification cs.CL
keywords grpolanguagemodelreasoningvisualmazemodelsspatial
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have demonstrated impressive capabilities in language processing, yet they often struggle with tasks requiring genuine visual spatial reasoning. In this paper, we introduce a novel two-stage training framework designed to equip standard LLMs with visual reasoning abilities for maze navigation. First, we leverage Supervised Fine Tuning (SFT) on a curated dataset of tokenized maze representations to teach the model to predict step-by-step movement commands. Next, we apply Group Relative Policy Optimization (GRPO)-a technique used in DeepSeekR1-with a carefully crafted reward function to refine the model's sequential decision-making and encourage emergent chain-of-thought behaviors. Experimental results on synthetically generated mazes show that while a baseline model fails to navigate the maze, the SFT-trained model achieves 86% accuracy, and further GRPO fine-tuning boosts accuracy to 93%. Qualitative analyses reveal that GRPO fosters more robust and self-corrective reasoning, highlighting the potential of our approach to bridge the gap between language models and visual spatial tasks. These findings offer promising implications for applications in robotics, autonomous navigation, and other domains that require integrated visual and sequential reasoning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OmniRouting: A Semantic-Coupled Multimodal Benchmark for Constraint-Aware Spatial Reasoning in PCB Routing

    cs.CV 2026-08 conditional novelty 7.0 of 10

    The OmniRouting benchmark, with 1,681 PCB designs, shows current large multimodal models achieve under 13% clean net routability while humans reach about 94%, exposing major gaps in constraint-aware spatial reasoning.

  2. Reasoning LLMs are Wandering Solution Explorers

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Six current reasoning LLMs, including commercial systems, exhibit structured-search failures on verifiable computation tasks and degrade as the solution space grows.

  3. Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding

    cs.CV 2025-11 reject novelty 5.0 of 10

    A new video benchmark and post-training recipe claim to improve VLMs' counterfactual 'what if' reasoning, but the reported gains likely come from training on the test set.

  4. Constructing coherent spatial memory in LLM agents through graph rectification

    cs.AI 2025-10 conditional novelty 5.0 of 10

    LLM-MapRepair uses versioned graph history and an edge-impact score to detect and repair structural errors in incrementally built LLM navigation graphs, improving repair accuracy from ~6% to ~55% on cleaned MANGO games.

  5. Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test

    cs.CL 2025-08 conditional novelty 5.0 of 10

    A new maze-navigation benchmark for LLMs reports that reasoning models outperform standard ones, but the link from performance gaps to a lack of persistent self-awareness is an overreach.

  6. Large Language Models for Planning: A Comprehensive and Systematic Survey

    cs.AI 2025-05 conditional novelty 3.0 of 10

    A structured survey of LLM planning methods, benchmarks, and interpretability work, organized around a three-way taxonomy.

Pith tools