Pith. sign in

REVIEW 10 cited by

Magma: A Foundation Model for Multimodal AI Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.13130 v1 pith:BKEN5DMQ submitted 2025-02-18 cs.CV cs.AIcs.HCcs.LGcs.RO

classification cs.CVcs.AIcs.HCcs.LGcs.RO
keywords magmatasksmodelmultimodalagenticintelligencemodelsability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Magma, a foundation model that serves multimodal AI agentic tasks in both the digital and physical worlds. Magma is a significant extension of vision-language (VL) models in that it not only retains the VL understanding ability (verbal intelligence) of the latter, but is also equipped with the ability to plan and act in the visual-spatial world (spatial-temporal intelligence) and complete agentic tasks ranging from UI navigation to robot manipulation. To endow the agentic capabilities, Magma is pretrained on large amounts of heterogeneous datasets spanning from images, videos to robotics data, where the actionable visual objects (e.g., clickable buttons in GUI) in images are labeled by Set-of-Mark (SoM) for action grounding, and the object movements (e.g., the trace of human hands or robotic arms) in videos are labeled by Trace-of-Mark (ToM) for action planning. Extensive experiments show that SoM and ToM reach great synergy and facilitate the acquisition of spatial-temporal intelligence for our Magma model, which is fundamental to a wide range of tasks as shown in Fig.1. In particular, Magma creates new state-of-the-art results on UI navigation and robotic manipulation tasks, outperforming previous models that are specifically tailored to these tasks. On image and video-related multimodal tasks, Magma also compares favorably to popular large multimodal models that are trained on much larger datasets. We make our model and code public for reproducibility at https://microsoft.github.io/Magma.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents

    cs.CL 2026-08 accept novelty 7.0 of 10

    Per-task routing between text, image, and hybrid observations of a browser page does not currently beat one fixed choice, because the labels needed to learn routing exist only where the agent already succeeds; only a ...

  2. RoboTTT: Context Scaling for Robot Policies

    cs.RO 2026-07 conditional novelty 7.0 of 10

    A robot policy that updates its own weights during deployment can use 8,000 steps of history, steadily improving as context grows and enabling one-shot imitation from human videos.

  3. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  4. RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.

  5. From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

    cs.RO 2025-06 conditional novelty 6.0 of 10

    INT-ACT is a new 50-task simulation benchmark showing that VLA robot policies maintain high intention accuracy but much lower grasp and task success under out-of-distribution conditions.

  6. GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents

    cs.CL 2025-06 conditional novelty 6.0 of 10

    An attention-based action head with multi-patch supervision outperforms coordinate-generation baselines on GUI grounding, and a verifier further improves accuracy.

  7. ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A GRPO-based reinforcement learning framework teaches an MLLM to propose zoom-in regions, improving small-object detection and interactive segmentation under a fixed sensing budget.

  8. Self++: Merging Human and AI for Co-Determined XR Living in the Metaverse

    cs.HC 2025-07 conditional novelty 5.0 of 10

    Self++ introduces a nine-level, SDT-based framework for co-determined human-AI living in extended reality, spanning competence, autonomy, and relatedness.

  9. 3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A diffusion world model predicts 3D optical flow as an embodiment-agnostic action plan, and constrained optimization converts the flow into robot arm actions.

  10. A Survey: Learning Embodied Intelligence from Physical Simulators and World Models

    cs.RO 2025-07 conditional novelty 4.0 of 10

    Embodied intelligence learning is reviewed through the complementary lenses of physical simulators and world models, with a proposed IR-L0 to IR-L4 robot capability taxonomy.

Pith tools