Pith. sign in

REVIEW 2 cited by

MUTEX: Learning Unified Policies from Multimodal Task Specifications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.14320 v1 pith:NZ6WGPCP submitted 2023-09-25 cs.RO

classification cs.RO
keywords mutextaskcross-modalgoalinstructionslearningmodalitiesspecifications
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans use different modalities, such as speech, text, images, videos, etc., to communicate their intent and goals with teammates. For robots to become better assistants, we aim to endow them with the ability to follow instructions and understand tasks specified by their human partners. Most robotic policy learning methods have focused on one single modality of task specification while ignoring the rich cross-modal information. We present MUTEX, a unified approach to policy learning from multimodal task specifications. It trains a transformer-based architecture to facilitate cross-modal reasoning, combining masked modeling and cross-modal matching objectives in a two-stage training procedure. After training, MUTEX can follow a task specification in any of the six learned modalities (video demonstrations, goal images, text goal descriptions, text instructions, speech goal descriptions, and speech instructions) or a combination of them. We systematically evaluate the benefits of MUTEX in a newly designed dataset with 100 tasks in simulation and 50 tasks in the real world, annotated with multiple instances of task specifications in different modalities, and observe improved performance over methods trained specifically for any single modality. More information at https://ut-austin-rpl.github.io/MUTEX/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition

    cs.RO 2025-06 conditional novelty 6.0 of 10

    CARMA combines object detection, person tracking, action detection, and a vision-language model to produce instance-level actor-action-object triplets for human-robot group interactions, achieving up to 72% task succe...

  2. Adaptive Perception for Unified Visual Multi-modal Object Tracking

    cs.CV 2025-02 conditional novelty 5.0 of 10

    APTrack shows that a single unified multi-modal tracker using equal modality modeling and learnable token interaction can beat both unified and task-specific trackers on RGB-T, RGB-D, and RGB-E benchmarks.

Pith tools