Pith. sign in

REVIEW 2 cited by

Towards Comprehensive Multimodal Perception: Introducing the Touch-Language-Vision Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.09813 v3 pith:XUPGNB5V submitted 2024-03-14 cs.CV cs.RO

classification cs.CVcs.RO
keywords touch-language-visionalignmentdatasetdescriptionsmultimodalpageperceptionsentence-level
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Tactility provides crucial support and enhancement for the perception and interaction capabilities of both humans and robots. Nevertheless, the multimodal research related to touch primarily focuses on visual and tactile modalities, with limited exploration in the domain of language. Beyond vocabulary, sentence-level descriptions contain richer semantics. Based on this, we construct a touch-language-vision dataset named TLV (Touch-Language-Vision) by human-machine cascade collaboration, featuring sentence-level descriptions for multimode alignment. The new dataset is used to fine-tune our proposed lightweight training framework, STLV-Align (Synergistic Touch-Language-Vision Alignment), achieving effective semantic alignment with minimal parameter adjustments (1%). Project Page: https://xiaoen0.github.io/touch.page/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals

    cs.CL 2025-07 conditional novelty 6.0 of 10

    HapticCap is the first large human-annotated vibration-caption dataset, and a contrastive retrieval model using T5 and AST achieves the best caption-matching performance among the tested baselines.

  2. RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual Data

    cs.CV 2025-05 conditional novelty 6.0 of 10

    RA-Touch uses tactile-guided retrieval over a GPT-recaptioned ImageNet dataset to improve open-vocabulary tactile description generation, raising TVL scores from 5.03 to 5.36.

Pith tools