REVIEW 2 cited by
Towards Comprehensive Multimodal Perception: Introducing the Touch-Language-Vision Dataset
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Tactility provides crucial support and enhancement for the perception and interaction capabilities of both humans and robots. Nevertheless, the multimodal research related to touch primarily focuses on visual and tactile modalities, with limited exploration in the domain of language. Beyond vocabulary, sentence-level descriptions contain richer semantics. Based on this, we construct a touch-language-vision dataset named TLV (Touch-Language-Vision) by human-machine cascade collaboration, featuring sentence-level descriptions for multimode alignment. The new dataset is used to fine-tune our proposed lightweight training framework, STLV-Align (Synergistic Touch-Language-Vision Alignment), achieving effective semantic alignment with minimal parameter adjustments (1%). Project Page: https://xiaoen0.github.io/touch.page/.
Forward citations
Cited by 2 Pith papers
-
HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals
HapticCap is the first large human-annotated vibration-caption dataset, and a contrastive retrieval model using T5 and AST achieves the best caption-matching performance among the tested baselines.
-
RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual Data
RA-Touch uses tactile-guided retrieval over a GPT-recaptioned ImageNet dataset to improve open-vocabulary tactile description generation, raising TVL scores from 5.03 to 5.36.
Discussion (0). Continue with ORCID to comment.