Pith. sign in

REVIEW 10 cited by

VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.09577 v1 pith:FEAOCOWR submitted 2025-05-14 cs.RO

classification cs.RO
keywords modelvtlaeffectivelyinsertionintroducelearningmanipulationmulti-modal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While vision-language models have advanced significantly, their application in language-conditioned robotic manipulation is still underexplored, especially for contact-rich tasks that extend beyond visually dominant pick-and-place scenarios. To bridge this gap, we introduce Vision-Tactile-Language-Action model, a novel framework that enables robust policy generation in contact-intensive scenarios by effectively integrating visual and tactile inputs through cross-modal language grounding. A low-cost, multi-modal dataset has been constructed in a simulation environment, containing vision-tactile-action-instruction pairs specifically designed for the fingertip insertion task. Furthermore, we introduce Direct Preference Optimization (DPO) to offer regression-like supervision for the VTLA model, effectively bridging the gap between classification-based next token prediction loss and continuous robotic tasks. Experimental results show that the VTLA model outperforms traditional imitation learning methods (e.g., diffusion policies) and existing multi-modal baselines (TLA/VLA), achieving over 90% success rates on unseen peg shapes. Finally, we conduct real-world peg-in-hole experiments to demonstrate the exceptional Sim2Real performance of the proposed VTLA model. For supplementary videos and results, please visit our project website: https://sites.google.com/view/vtla

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich Manipulation

    cs.RO 2026-07 conditional novelty 7.0 of 10

    A vision-language-action robot policy that stores compressed wrist-force histories as memory tokens can count contact events that are visually ambiguous, achieving 83.3% average success across three tasks.

  2. Feeling the Unexpected: ResTacVLA for Contact-Rich Manipulation via Residual Tactile Representation

    cs.RO 2026-07 conditional novelty 6.5 of 10

    Residual tactile representations plus surprise-aware gating let VLA policies master contact-rich robot tasks that vision-only models fail.

  3. Representation-Aligned Tactile Grounding for Contact-Rich Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Future tactile prediction applied to intermediate action-expert features, rather than visual-language or final-action features, improves contact-rich manipulation in SmolVLA and π0.

  4. SoftVTBench: A Safety-Aware Visuo-Tactile Benchmark for Physically Constrained Robotic Manipulation of Deformable Objects

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Success-only metrics overstate deformable-manipulation performance; tactile sensing raises Safety Success (e.g. 21.4%→35.6% on Object-Soft) while Goal Success stays comparable.

  5. Tactile Modality Fusion for Vision-Language-Action Models

    cs.RO 2026-03 conditional novelty 6.0 of 10

    A FiLM-based tactile fusion method that conditions VLA visual features on frozen pretrained touch embeddings improves real-robot insertion success, speed, and force control relative to vision-only and concatenation baselines.

  6. EquiBim: Learning Symmetry-Equivariant Policy for Bimanual Manipulation

    cs.RO 2026-03 conditional novelty 6.0 of 10

    Adding a loss that enforces left-right equivariance between observations and actions improves average bimanual imitation policy success by +2.7 to +9.5 points across four observation/action settings.

  7. Force-Aware Residual DAgger via Trajectory Editing for Precision Insertion with Impedance Control

    cs.RO 2026-03 conditional novelty 6.0 of 10

    TER-DAgger uses force-prediction mismatches to trigger human corrections and residual-policy training, lifting precision-insertion success from 40.0% to 77.2% on average.

  8. DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter

    cs.RO 2026-02 conditional novelty 6.0 of 10

    A decoupled multimodal diffusion transformer with a LoRA tactile adapter improves real-world bimanual manipulation success by 21 percentage points over a diffusion-policy baseline; a new 50-hour tactile bimanual datas...

  9. TacWAM: Anchor-Guided World Action Model with Mechanics-Aware Tactile Prediction

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Mechanics-aware future tactile prediction plus history and train-time isolation of future tokens raises real-robot contact-rich success from ~37.5% to 75% average.

  10. HapticVLA: Contact-Rich Manipulation via Vision-Language-Action Model without Inference-Time Tactile Sensing

    cs.RO 2026-03 conditional novelty 5.0 of 10

    A vision-language-action robot policy that claims tactile-aware manipulation without tactile sensors at inference via reward-weighted flow matching and action distillation.

Pith tools