Pith. sign in

REVIEW 6 cited by

Self-Supervised Visuo-Tactile Pretraining to Locate and Follow Garment Features

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.13042 v2 pith:2I6UQJHP submitted 2022-09-26 cs.RO

classification cs.RO
keywords featurerepresentationsself-supervisedtactilevisualvisuo-tactilecross-modalgarment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans make extensive use of vision and touch as complementary senses, with vision providing global information about the scene and touch measuring local information during manipulation without suffering from occlusions. While prior work demonstrates the efficacy of tactile sensing for precise manipulation of deformables, they typically rely on supervised, human-labeled datasets. We propose Self-Supervised Visuo-Tactile Pretraining (SSVTP), a framework for learning multi-task visuo-tactile representations in a self-supervised manner through cross-modal supervision. We design a mechanism that enables a robot to autonomously collect precisely spatially-aligned visual and tactile image pairs, then train visual and tactile encoders to embed these pairs into a shared latent space using cross-modal contrastive loss. We apply this latent space to downstream perception and control of deformable garments on flat surfaces, and evaluate the flexibility of the learned representations without fine-tuning on 5 tasks: feature classification, contact localization, anomaly detection, feature search from a visual query (e.g., garment feature localization under occlusion), and edge following along cloth edges. The pretrained representations achieve a 73-100% success rate on these 5 tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. {\tau}: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Action-conditioned JEPA-style future-visual latent prediction yields dynamics-aware tactile tokens that lift contact-rich VLA success rates from ~30% to ~70% average on four real tasks.

  2. Tactile MNIST: Benchmarking Active Tactile Perception

    cs.RO 2025-06 conditional novelty 6.0 of 10

    The authors release a Gymnasium-compatible benchmark with four active tactile tasks, 13,580 3D digit models, and 153,600 real touches on 600 printed digits.

  3. Universal Visuo-Tactile Video Understanding for Embodied Interaction

    cs.CV 2025-05 conditional novelty 6.0 of 10

    VTV-LLM is a tactile-video large language model, trained on a new VTV150K dataset, that reasons about hardness, protrusion, elasticity, and friction in natural language.

  4. RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual Data

    cs.CV 2025-05 conditional novelty 6.0 of 10

    RA-Touch uses tactile-guided retrieval over a GPT-recaptioned ImageNet dataset to improve open-vocabulary tactile description generation, raising TVL scores from 5.03 to 5.36.

  5. Adaptive Visuo-Tactile Fusion with Predictive Force Attention for Dexterous Manipulation

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A force-guided attention module and future-force prediction auxiliary task improve visuo-tactile fusion for dexterous manipulation, reaching 93% average success in real robot trials.

  6. ConViTac: Aligning Visual-Tactile Fusion with Contrastive Representations

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A visual-tactile fusion network that conditions cross-modal attention on SimCLR contrastive embeddings improves material classification and grasp-success prediction in real-robot datasets.

Pith tools