Pith. sign in

REVIEW 10 cited by

Octopi: Object Property Reasoning with Large Tactile-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.02794 v2 pith:FYJKCFVE submitted 2024-05-05 cs.RO

classification cs.RO
keywords physicalreasoninglanguageoctopitactilephysiclearpropertytasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Physical reasoning is important for effective robot manipulation. Recent work has investigated both vision and language modalities for physical reasoning; vision can reveal information about objects in the environment and language serves as an abstraction and communication medium for additional context. Although these works have demonstrated success on a variety of physical reasoning tasks, they are limited to physical properties that can be inferred from visual or language inputs. In this work, we investigate combining tactile perception with language, which enables embodied systems to obtain physical properties through interaction and apply commonsense reasoning. We contribute a new dataset PhysiCLeAR, which comprises both physical/property reasoning tasks and annotated tactile videos obtained using a GelSight tactile sensor. We then introduce Octopi, a system that leverages both tactile representation learning and large vision-language models to predict and reason about tactile inputs with minimal language fine-tuning. Our evaluations on PhysiCLeAR show that Octopi is able to effectively use intermediate physical property predictions to improve its performance on various tactile-related tasks. PhysiCLeAR and Octopi are available at https://github.com/clear-nus/octopi.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TacReasoner: A Dynamic Tactile-Language Framework for Interactive Reasoning in Real-World Scenarios

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A dynamic-aware tactile encoder plus TouchCoT-10k chain-of-thought data lets a 7B model outperform larger tactile-language baselines on physical-property and real-world reasoning tasks.

  2. TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A tactile-aware world model recognizes failure-adjacent contact states, imagines local visuo-tactile corrections, and post-trains VLAs with knowledge insulation, raising average success by 44% over the base policy.

  3. Tactile Modality Fusion for Vision-Language-Action Models

    cs.RO 2026-03 conditional novelty 6.0 of 10

    A FiLM-based tactile fusion method that conditions VLA visual features on frozen pretrained touch embeddings improves real-robot insertion success, speed, and force control relative to vision-only and concatenation baselines.

  4. EmbRACE-3K: Embodied Reasoning and Action in Complex Environments

    cs.CV 2025-07 conditional novelty 6.0 of 10

    EmbRACE-3K provides a photorealistic, closed-loop embodied benchmark with stepwise reasoning annotations, and fine-tuning Qwen2.5-VL on it improves embodied task success in simulation.

  5. Tactile MNIST: Benchmarking Active Tactile Perception

    cs.RO 2025-06 conditional novelty 6.0 of 10

    The authors release a Gymnasium-compatible benchmark with four active tactile tasks, 13,580 3D digit models, and 153,600 real touches on 600 printed digits.

  6. Universal Visuo-Tactile Video Understanding for Embodied Interaction

    cs.CV 2025-05 conditional novelty 6.0 of 10

    VTV-LLM is a tactile-video large language model, trained on a new VTV150K dataset, that reasons about hardness, protrusion, elasticity, and friction in natural language.

  7. RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual Data

    cs.CV 2025-05 conditional novelty 6.0 of 10

    RA-Touch uses tactile-guided retrieval over a GPT-recaptioned ImageNet dataset to improve open-vocabulary tactile description generation, raising TVL scores from 5.03 to 5.36.

  8. SiPhy: Single-Image Physical Property Reasoning

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A single-image vision-language pipeline reports state-of-the-art mass, density, and stiffness predictions by combining CLIP features, a fine-tuned VLM, and depth-adaptive pseudo-voxel sampling.

  9. Robotic Perception with a Large Tactile-Vision-Language Model for Physical Property Inference

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A tactile-vision-language model predicts hardness, elasticity, and roughness with Spearman correlations up to 0.64, outperforming unimodal baselines on 35 objects.

  10. Foundation Model Driven Robotics: A Comprehensive Review

    cs.RO 2025-07 conditional novelty 2.0 of 10

    A review of foundation-model-driven robotics that synthesizes recent work across perception, planning, control, HRI, simulation, and sim-to-real transfer, and highlights open challenges.

Pith tools