Pith. sign in

REVIEW 3 cited by

Multimodal Visual-Tactile Representation Learning through Self-Supervised Contrastive Pre-Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.12024 v1 pith:OSEF2B33 submitted 2024-01-22 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords learningmvitacself-supervisedsensoryclassificationcontrastivegraspingleverages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rapidly evolving field of robotics necessitates methods that can facilitate the fusion of multiple modalities. Specifically, when it comes to interacting with tangible objects, effectively combining visual and tactile sensory data is key to understanding and navigating the complex dynamics of the physical world, enabling a more nuanced and adaptable response to changing environments. Nevertheless, much of the earlier work in merging these two sensory modalities has relied on supervised methods utilizing datasets labeled by humans.This paper introduces MViTac, a novel methodology that leverages contrastive learning to integrate vision and touch sensations in a self-supervised fashion. By availing both sensory inputs, MViTac leverages intra and inter-modality losses for learning representations, resulting in enhanced material property classification and more adept grasping prediction. Through a series of experiments, we showcase the effectiveness of our method and its superiority over existing state-of-the-art self-supervised and supervised techniques. In evaluating our methodology, we focus on two distinct tasks: material classification and grasping success prediction. Our results indicate that MViTac facilitates the development of improved modality encoders, yielding more robust representations as evidenced by linear probing assessments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning

    cs.RO 2025-11 unverdicted novelty 6.0 of 10

    MSDP pre-trains a transformer encoder with masked multisensory autoencoding, then uses an asymmetric actor-critic bridge (cross-attention for critic, pooling for actor) to accelerate and robustify contact-rich RL acro...

  2. AdaDexGrasp: Adaptive Dexterous Grasping via 3D Visuo-Tactile Representation Fusion

    cs.RO 2026-08 conditional novelty 5.0 of 10

    AdaDexGrasp learns to fuse point clouds with finger-level tactile labels to generate, judge, and correct dexterous grasps, reporting 91%/82%/83% success on seen, unseen-object, and unseen-category sets in simulation.

  3. ConViTac: Aligning Visual-Tactile Fusion with Contrastive Representations

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A visual-tactile fusion network that conditions cross-modal attention on SimCLR contrastive embeddings improves material classification and grasp-success prediction in real-robot datasets.

Pith tools