Pith. sign in

REVIEW 2 cited by

Touch2Touch: Cross-Modal Tactile Generation for Object Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.08269 v1 pith:MDHS56K4 submitted 2024-09-12 cs.RO

classification cs.RO
keywords sensorssensortouchbubblecross-modalgelslimmethodsmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Today's touch sensors come in many shapes and sizes. This has made it challenging to develop general-purpose touch processing methods since models are generally tied to one specific sensor design. We address this problem by performing cross-modal prediction between touch sensors: given the tactile signal from one sensor, we use a generative model to estimate how the same physical contact would be perceived by another sensor. This allows us to apply sensor-specific methods to the generated signal. We implement this idea by training a diffusion model to translate between the popular GelSlim and Soft Bubble sensors. As a downstream task, we perform in-hand object pose estimation using GelSlim sensors while using an algorithm that operates only on Soft Bubble signals. The dataset, the code, and additional details can be found at https://www.mmintlab.com/research/touch2touch/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. 3D Cal: An Open-Source Software Library for Depth Reconstruction on Vision-Based Tactile Sensors

    cs.RO 2025-11 conditional novelty 6.0 of 10

    3D Cal repurposes a 3D printer as an automated calibration rig and trains a lightweight CNN, TouchNet, to reconstruct depth maps for DIGIT and GelSight Mini.

  2. Universal Visuo-Tactile Video Understanding for Embodied Interaction

    cs.CV 2025-05 conditional novelty 6.0 of 10

    VTV-LLM is a tactile-video large language model, trained on a new VTV150K dataset, that reasons about hardness, protrusion, elasticity, and friction in natural language.

Pith tools