REVIEW 6 cited by
Self-Supervised Visuo-Tactile Pretraining to Locate and Follow Garment Features
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Humans make extensive use of vision and touch as complementary senses, with vision providing global information about the scene and touch measuring local information during manipulation without suffering from occlusions. While prior work demonstrates the efficacy of tactile sensing for precise manipulation of deformables, they typically rely on supervised, human-labeled datasets. We propose Self-Supervised Visuo-Tactile Pretraining (SSVTP), a framework for learning multi-task visuo-tactile representations in a self-supervised manner through cross-modal supervision. We design a mechanism that enables a robot to autonomously collect precisely spatially-aligned visual and tactile image pairs, then train visual and tactile encoders to embed these pairs into a shared latent space using cross-modal contrastive loss. We apply this latent space to downstream perception and control of deformable garments on flat surfaces, and evaluate the flexibility of the learned representations without fine-tuning on 5 tasks: feature classification, contact localization, anomaly detection, feature search from a visual query (e.g., garment feature localization under occlusion), and edge following along cloth edges. The pretrained representations achieve a 73-100% success rate on these 5 tasks.
Forward citations
Cited by 6 Pith papers
-
{\tau}: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision
Action-conditioned JEPA-style future-visual latent prediction yields dynamics-aware tactile tokens that lift contact-rich VLA success rates from ~30% to ~70% average on four real tasks.
-
Tactile MNIST: Benchmarking Active Tactile Perception
The authors release a Gymnasium-compatible benchmark with four active tactile tasks, 13,580 3D digit models, and 153,600 real touches on 600 printed digits.
-
Universal Visuo-Tactile Video Understanding for Embodied Interaction
VTV-LLM is a tactile-video large language model, trained on a new VTV150K dataset, that reasons about hardness, protrusion, elasticity, and friction in natural language.
-
RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual Data
RA-Touch uses tactile-guided retrieval over a GPT-recaptioned ImageNet dataset to improve open-vocabulary tactile description generation, raising TVL scores from 5.03 to 5.36.
-
Adaptive Visuo-Tactile Fusion with Predictive Force Attention for Dexterous Manipulation
A force-guided attention module and future-force prediction auxiliary task improve visuo-tactile fusion for dexterous manipulation, reaching 93% average success in real robot trials.
-
ConViTac: Aligning Visual-Tactile Fusion with Contrastive Representations
A visual-tactile fusion network that conditions cross-modal attention on SimCLR contrastive embeddings improves material classification and grasp-success prediction in real-robot datasets.
Discussion (0). Continue with ORCID to comment.