REVIEW 2 major objections 5 minor 4 cited by
Tactile Beyond Pixels: Multisensory Touch Representations for Robot Manipulation
T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A self-supervised transformer, pretrained on roughly one million unlabeled multisensory touch interactions from the Digit 360 sensor, produces representations that outperform end-to-end tactile-image training on manipulation policies and…
desk verdict Solid empirical core, useful architecture—but the flagship 48% benchmark claim in the abstract is not backed by the reported per-task numbers, and the object-action benchmark is in-distribution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is Sparsh-X itself: a 12-layer transformer (8 unimodal layers followed by 4 fusion layers) that fuses four tactile modalities through four learnable bottleneck tokens shared across modalities. Each modality is tokenized separately—image patches, log-mel audio spectrograms, and temporal chunks of IMU and pressure—and processed independently before cross-modal attention. Pretraining uses a teacher–student self-distillation objective with per-modality masking, where teacher tokens serve as pseudo-labels; this is what makes the representations self-supervised and scalable. The other named mechanism is tactile adaptation via ControlNet, which attaches a zero-initialized convolutional path to a frozen simulation-trained policy (Hora) so that Sparsh-X features can correct the policy's privileged latent without retraining the base model.
What would settle it
Take the same Sparsh-X training recipe but pretrain on data from a single Digit 360 sensor and a single surface; if downstream gains over end-to-end training disappear, then the reported 48% average improvement is attributable to data diversity rather than to multisensory fusion or self-supervised pretraining. Conversely, evaluating Sparsh-X on a genuinely held-out object–action–surface combination that was never seen in pretraining, and comparing against an end-to-end model trained on that task, would directly test whether the representation transfers beyond the training distribution.
Extended reading notes
Core claim
The central claim is that a single general-purpose backbone, pretrained without labels on roughly one million multisensory touch interactions, can represent contact in a way that transfers across policies and inference tasks. Sparsh-X takes Digit 360 signals—30 fps tactile images, contact-microphone audio, accelerometer motion, and static pressure—and maps them into a shared latent space through a transformer that first processes each modality independently and then fuses them with bottleneck attention tokens. When frozen and fed to downstream decoders, these latents improve object-action-surface classification, material-quantity estimation, and normal-force regression over end-to-end models trained from scratch, and they improve real-world plug insertion and in-hand rotation policies, including tactile adaptation of a simulation-trained policy via a zero-convolution ControlNet module. The authors present Sparsh-X as the first multisensory touch representation and position it as a step toward foundation models for touch.
Load-bearing premise
The pretraining dataset—under a million samples from just six Digit 360 sensors, gathered by an Allegro hand rummaging in a tray and a handheld picker tapping and sliding over four surfaces—must be sufficiently diverse that representations learned from it generalize to the downstream tasks; the main classification benchmark is evaluated on a subset of that same data.
Editorial extensions
If this is right
- Frozen Sparsh-X representations, fed through an attentive pooling head, lift plug-insertion success to 90% over 20 trials, a 63% gain over an end-to-end tactile-image policy and 500% over external-vision-only.
- Tactile adaptation with Sparsh-X cuts vertical drift in in-hand rotation by 90% relative to the Hora baseline, and all-modality inputs keep the object stable when friction is reduced, where image-only features fail.
- Multisensory fusion plus self-supervised pretraining delivers roughly 48% average accuracy improvement over end-to-end tactile-image training across object-action-surface, material-quantity, and force-estimation benchmarks, with larger margins under low data budgets.
- Because Sparsh-X runs at about 20 Hz on four Digit 360 fingertips with a single RTX 4090, the representation is deployable in closed-loop real-world policies as used in the insertion and rotation experiments.
Reading between the lines
- If Sparsh-X-style representations prove transferable across sensor hardware and embodiments, they could become a standard input layer for contact-rich policies, letting sim-to-real transfer consume learned touch latents instead of raw high-dimensional signals; the paper only demonstrates this on one sensor family.
- The paper explicitly leaves fine-tuning untested; since its own insertion experiment shows end-to-end image encoders beating frozen pretrained image features, fine-tuning Sparsh-X on task data is a likely lever to close that gap and may also fix the image-modality diversity weakness the authors note.
- The bottleneck-fusion recipe suggests a scaling path: if pretraining data diversity grows (more sensors, objects, surfaces), representation quality may continue to improve without architectural changes—this is a testable prediction, not something the paper demonstrates.
- The 63% and 48% gains are in-distribution or near-distribution; the strongest open question the paper leaves for the field is whether multisensory touch representations generalize to contact scenarios far outside the six-sensor, two-platform training set.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Sparsh-X, a transformer-based backbone that fuses four tactile modalities from the Digit 360 sensor (image, audio, IMU/accelerometer, and pressure) and is pretrained with self-supervised learning on roughly one million unlabeled contact interactions. The learned representations are frozen and evaluated on three supervised physical-property tasks (object-action-surface classification, material-quantity estimation, normal-force regression) and on two real-robot policy tasks (plug insertion via imitation learning, and in-hand rotation via a proposed ControlNet-style tactile adaptation of a sim-trained policy). The central claims are that multisensory pretraining improves physical-property inference by 48% on average over end-to-end tactile-image-only baselines, boosts plug-insertion success by 63%, and reduces vertical drift in in-hand rotation by 90%.
Significance. If the empirical claims hold, Sparsh-X would be a useful contribution to robot touch representation learning: it addresses a real gap by moving beyond image-only tactile sensing to a unified multisensory representation, and it evaluates on real hardware with ablations over input modalities and training-data budgets. The use of frozen representations to isolate pretraining quality, the inclusion of data-efficiency curves, and the deployment of tactile adaptation on top of an open-sourced sim policy (Hora) are concrete strengths. The paper does not provide machine-checked proofs or parameter-free derivations, but its contribution is empirical and the experimental scope is appropriate for a robotics venue. The two load-bearing issues identified below, however, affect the paper's headline quantitative claims and must be resolved before the results can be taken at face value.
major comments (2)
- [Abstract and §5] The abstract and §5 state that Sparsh-X achieves an average improvement of 48% over end-to-end tactile-image-only models across the physical-property tasks. The only per-task numbers reported in §4.1 are 13% for object-action-surface classification, 20.5% for material-quantity estimation, and 17% for normal-force estimation. These three numbers average to about 17%, and their sum is about 50%, not 48%. Moreover, the 13% figure is described as a gain within Sparsh-X (all modalities versus image only), not as a comparison against an end-to-end image-only baseline, so the three numbers are not obviously commensurable. No table or appendix entry gives per-task values that average to 48%. Please correct the abstract and §5 to report the actual mean, or provide the explicit aggregation rule, baseline definitions, and data budgets that yield 48%.
- [§4.1 and Appendix A.2] The object-action-surface classification benchmark is evaluated on a re-annotated subset of the very same manual-picker sequences used for pretraining: Appendix A.2 states that the authors 'repurpose the SSL pretraining dataset collected with the manual picker' and that the 104 available sequences are split into 69 training and 35 testing. The frozen Sparsh-X encoder has therefore already seen this distribution, including the same Digit 360 sensors, objects, actions, and surfaces, whereas the end-to-end baseline sees this data only during classifier training. The 13% gain and the confusion matrix in Figure 12 are thus an in-distribution comparison and do not by themselves demonstrate transfer or generalization of the representation. The material-quantity and normal-force tasks, which use separate data collections, do support cross-task transfer, but the object-action-surface result and its role in the claimed 48% average should be explicitly labeled as in-distribution, or the evaluation should be repeated on held-out objects, surfaces, or sensor instances.
minor comments (5)
- [Appendix B and Figure 14 caption] The manuscript refers to 'TacX' in two places ('the TacX-based classifier' in Appendix B and 'how the inputs are constructed for TacX' in the Figure 14 caption), but the model is called Sparsh-X throughout the rest of the paper. Please replace these leftover names.
- [Section 3 and Appendix A.1] The temporal window lengths are inconsistent: Section 3 says audio and IMU use 0.55s windows and pressure uses a 1.1s window, while Appendix A.1 says 0.5s for audio/IMU and 1s for pressure, and Figure 2 labels the pressure window as 1.0s. Please align these numbers.
- [Introduction] The first paragraph of the Introduction is duplicated verbatim; the second copy should be removed.
- [Throughout] The paper alternates between 'motion' and 'IMU/accelerometer' for the same modality, and between 'pressure' and 'static pressure sensor.' Please define the modality names once and use them consistently.
- [Figures 4, 5, and 7] The bar charts report only point estimates with no error bars, confidence intervals, or per-trial outcome tables. Given that the 63%, 90%, and 48% claims are central, please add trial counts, per-condition values, and preferably confidence intervals or raw outcomes for the policy evaluations and the main benchmark comparisons.
Circularity Check
No circularity: the reported numbers are empirical comparisons, not derivations from fitted inputs or a self-citation chain; the in-distribution benchmark overlap and the 48% arithmetic inconsistency are correctness limitations, not circular steps.
full rationale
The paper's central claims are measured outcomes. Sparsh-X is trained with a self-supervised objective (teacher-student self-distillation with masked token prediction) that does not use downstream labels, and the downstream policy and benchmark numbers (63% insertion improvement, 90% vertical-drift reduction, per-task accuracies) are empirical comparisons against end-to-end baselines. No parameter is fitted to a subset of data and then renamed as a prediction, and no load-bearing argument reduces to a self-citation. The only circularity-adjacent issue is that the object-action-surface benchmark repurposes the manual-picker portion of the pretraining corpus (Appendix A.2), so that evaluation is in-distribution and may overstate generalization; however, this is a data-split and leakage limitation rather than a logical reduction, because the pretraining targets are masked reconstruction/clustering pseudo-labels, not the classification labels used downstream. The paper also cites its own prior Sparsh work [9] for the 30 fps image sampling and the normal-force protocol, but these are minor technical conventions and not load-bearing. The abstract's '48% average improvement' figure is not reproduced by the three per-task margins reported in Section 4.1 (13%, 20.5%, and 17%, whose mean is about 17%), which is an arithmetic or aggregation inconsistency and a legitimate correctness concern, but not a circular derivation. Consequently, no circular step can be exhibited, and the circularity score is 0.
Assumptions & free parameters
free parameters (3)
- Architecture hyperparameters =
L=12, Lf=8, Lb=4, B=4, patch size 16, image resolution 224x224
- Masking ratios =
local masks retain 10-50%, global masks retain 50-100%
- Temporal windows per modality =
image 0.17s (two frames at 30fps with stride 5), audio 0.55s, IMU 0.55s, pressure 1.1s
assumptions (5)
- standard math Transformer attention and the DINOv2 self-distillation objective behave as described in the cited literature.
- domain assumption The Digit 360 sensor provides synchronized, complementary measurements across image, audio, IMU, and pressure.
- domain assumption Self-supervised pretraining on unlabeled contact data yields representations useful for downstream physical property inference.
- domain assumption The privileged information used by Hora can be approximated from tactile inputs via the ControlNet module without retraining the base policy.
- domain assumption The data collection protocols produce contact interactions representative of the downstream tasks.
Cite this review
Pith. "Pith review of Tactile Beyond Pixels: Multisensory Touch Representations for Robot Manipulation." pith.science (2026). https://pith.science/paper/VCABCPFV
@misc{pith2026250614754,
author = {Pith},
title = {Pith review of: Tactile Beyond Pixels: Multisensory Touch Representations for Robot Manipulation},
year = {2026},
howpublished = {\url{https://pith.science/paper/VCABCPFV}},
note = {Machine review of arXiv:2506.14754}
}
read the original abstract
We present Sparsh-X, the first multisensory touch representations across four tactile modalities: image, audio, motion, and pressure. Trained on ~1M contact-rich interactions collected with the Digit 360 sensor, Sparsh-X captures complementary touch signals at diverse temporal and spatial scales. By leveraging self-supervised learning, Sparsh-X fuses these modalities into a unified representation that captures physical properties useful for robot manipulation tasks. We study how to effectively integrate real-world touch representations for both imitation learning and tactile adaptation of sim-trained policies, showing that Sparsh-X boosts policy success rates by 63% over an end-to-end model using tactile images and improves robustness by 90% in recovering object states from touch. Finally, we benchmark Sparsh-X ability to make inferences about physical properties, such as object-action identification, material-quantity estimation, and force estimation. Sparsh-X improves accuracy in characterizing physical properties by 48% compared to end-to-end approaches, demonstrating the advantages of multisensory pretraining for capturing features essential for dexterous manipulation.
Figures
Figures from the paper (12 more)
Forward citations
Cited by 4 Pith papers
-
Tactile Genesis: Exploring Tactile Sensors at Scale for Learning Dexterous Tasks
Whole-hand tactile coverage and per-taxel force/torque dominate sensor type and resolution for learning three dexterous tasks in a new high-throughput tactile simulator.
-
Representation-Aligned Tactile Grounding for Contact-Rich Robotic Manipulation
Future tactile prediction applied to intermediate action-expert features, rather than visual-language or final-action features, improves contact-rich manipulation in SmolVLA and π0.
-
Online World Modeling Enables Real-World Inverse Reinforcement Learning from Observation
MPAIL2 demonstrates real-world manipulation learning from observation alone, without rewards or action labels, plus positive online transfer.
-
DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter
A decoupled multimodal diffusion transformer with a LoRA tactile adapter improves real-world bimanual manipulation success by 21 percentage points over a diffusion-policy baseline; a new 50-hour tactile bimanual datas...
Reference graph
Works this paper leans on
-
[1]
W. Yuan, S. Dong, and E. H. Adelson. GelSight: High-resolution robot tactile sensors for estimating geometry and force.Sensors, 17(12), 2017. ISSN 1424-8220. doi:10.3390/ s17122762. URLhttps://www.mdpi.com/1424-8220/17/12/2762
work page 2017
-
[2]
M. Lambeta, P.-W. Chou, S. Tian, B. Yang, B. Maloon, V . R. Most, D. Stroud, R. Santos, A. Byagowi, G. Kammerer, D. Jayaraman, and R. Calandra. Digit: A novel design for a low-cost compact high-resolution tactile sensor with application to in-hand manipulation.IEEE Robotics and Automation Letters, 5(3):3838–3845, 2020. doi:10.1109/LRA.2020.2977257
-
[3]
N. F. Lepora. Soft biomimetic optical tactile sensing with the tactip: A review.IEEE Sensors Journal, 21(19):21131–21143, 2021
2021
-
[4]
M. Lambeta, T. Wu, A. Sengul, V . R. Most, N. Black, K. Sawyer, R. Mercado, H. Qi, A. Sohn, B. Taylor, N. Tydingco, G. Kammerer, D. Stroud, J. Khatha, K. Jenkins, K. Most, N. Stein, R. Chavira, T. Craven-Bartle, E. Sanchez, Y . Ding, J. Malik, and R. Calandra. Digitizing touch with an artificial multimodal fingertip, 2024. URLhttps://arxiv.org/abs/2411.02479
arXiv 2024
-
[5]
K. Yu, Y . Han, Q. Wang, V . Saxena, D. Xu, and Y . Zhao. Mimictouch: Leveraging multi-modal human tactile demonstrations for contact-rich manipulation. In8th Annual Conference on Robot Learning, 2024. URLhttps://openreview.net/forum?id=7yMZAUkXa4
work page 2024
- [6]
-
[7]
Z. Xu, R. Uppuluri, X. Zhang, C. Fitch, P. G. Crandall, W. Shou, D. Wang, and Y . She. UniT: Data efficient tactile representation with generalization to unseen objects, 2025. URL https://arxiv.org/abs/2408.06481
arXiv 2025
-
[8]
J. Zhao, Y . Ma, L. Wang, and E. Adelson. Transferable tactile transformers for representation learning across diverse sensors and tasks. In8th Annual Conference on Robot Learning, 2024. URLhttps://openreview.net/forum?id=KXsropnmNI. 9
work page 2024
Show all 56 references
-
[9]
Higuera, A
C. Higuera, A. Sharma, C. K. Bodduluri, T. Fan, P. Lancaster, M. Kalakrishnan, M. Kaess, B. Boots, M. Lambeta, T. Wu, and M. Mukadam. Sparsh: Self-supervised touch representations for vision-based tactile sensing. In8th Annual Conference on Robot Learning, 2024. URL https://op...
2024
-
[10]
Gupta, Y
H. Gupta, Y . Mo, S. Jin, and W. Yuan. Sensor-invariant tactile representation.arXiv preprint arXiv:2502.19638, 2025
2025 arXiv
-
[11]
Oquab, T
M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al. DINOv2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023
2023 arXiv
-
[12]
Bardes, Q
A. Bardes, Q. Garrido, J. Ponce, X. Chen, M. Rabbat, Y . LeCun, M. Assran, and N. Ballas. V-JEPA: Latent video prediction for visual representation learning. 2023
2023
-
[13]
Hendrycks, M
D. Hendrycks, M. Mazeika, S. Kadavath, and D. Song. Using self-supervised learning can improve model robustness and uncertainty.Advances in neural information processing systems, 32, 2019
2019
-
[14]
Karnan, E
H. Karnan, E. Yang, D. Farkash, G. Warnell, J. Biswas, and P. Stone. Sterling: Self- supervised terrain representation learning from unconstrained robot experience.arXiv preprint arXiv:2309.15302, 2023
2023 arXiv
-
[15]
Y . Lin, J. Lloyd, A. Church, and N. F. Lepora. Tactile gym 2.0: Sim-to-real deep reinforcement learning for comparing low-cost high-resolution robot touch.IEEE Robotics and Automation Letters, 7(4):10754–10761, 2022
2022
-
[16]
Donlon, S
E. Donlon, S. Dong, M. Liu, J. Li, E. Adelson, and A. Rodriguez. Gelslim: A high- resolution, compact, robust, and calibrated tactile-sensing finger. In2018 IEEE/RSJ Inter- national Conference on Intelligent Robots and Systems (IROS), pages 1927–1934, 2018. doi: 10.1109/IROS.2...
1927
-
[17]
W. Yuan, Y . Mo, S. Wang, and E. H. Adelson. Active clothing material perception using tactile sensing and deep learning. In2018 IEEE International Conference on Robotics and Automation (ICRA), pages 4842–4849, 2018. doi:10.1109/ICRA.2018.8461164
2018
-
[18]
Guo, H.-J
X. Guo, H.-J. Huang, and W. Yuan. Estimating properties of solid particles inside container using touch sensing. In2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 8985–8992, 2023. doi:10.1109/IROS55552.2023.10341880
2023
-
[19]
Suresh, H
S. Suresh, H. Qi, T. Wu, T. Fan, L. Pineda, M. Lambeta, J. Malik, M. Kalakrishnan, R. Calandra, M. Kaess, et al. Neuralfeels with neural fields: Visuotactile perception for in-hand manipulation. Science Robotics, 9(96):eadl0628, 2024
2024
-
[20]
R. Gao, K. Deng, G. Yang, W. Yuan, and J.-Y . Zhu. Tactile dreamfusion: Exploiting tactile sensing for 3d generation. InConference on Neural Information Processing Systems (NeurIPS), 2024
2024
-
[21]
Suresh, Z
S. Suresh, Z. Si, S. Anderson, M. Kaess, and M. Mukadam. Midastouch: Monte-carlo inference over distributions across sliding touch. InConference on Robot Learning, pages 319–331. PMLR, 2023
2023
-
[22]
Huang, M
H.-J. Huang, M. Kaess, and W. Yuan. Normalflow: Fast, robust, and accurate contact-based object 6dof pose tracking with vision-based tactile sensors.IEEE Robotics and Automation Letters, 2024
2024
-
[23]
S. Dong, D. K. Jha, D. Romeres, S. Kim, D. Nikovski, and A. Rodriguez. Tactile-rl for insertion: Generalization to objects of unknown geometry. In2021 IEEE International Conference on Robotics and Automation (ICRA), pages 6437–6443. IEEE, 2021. 10
2021
-
[24]
Aquilina, D
K. Aquilina, D. A. Barton, and N. F. Lepora. Tactile control for object tracking and dynamic contour following.Robotics and Autonomous Systems, 178:104710, 2024
2024
-
[25]
H. Xue, J. Ren, W. Chen, G. Zhang, Y . Fang, G. Gu, H. Xu, and C. Lu. Reactive diffusion policy: Slow-fast visual-tactile policy learning for contact-rich manipulation.arXiv preprint arXiv:2503.02881, 2025
2025 arXiv
-
[26]
Gandhi, A
D. Gandhi, A. Gupta, and L. Pinto. Swoosh! rattle! thump!–actions that sound.arXiv preprint arXiv:2007.01851, 2020
2007 arXiv
-
[27]
Clarke, N
S. Clarke, N. Heravi, M. Rau, R. Gao, J. Wu, D. James, and J. Bohg. Diffimpact: Differentiable rendering and identification of impact sounds. InConference on Robot Learning, pages 662–673. PMLR, 2022
2022
-
[28]
Thankaraj and L
A. Thankaraj and L. Pinto. That sounds right: Auditory self-supervision for dynamic robot manipulation. InConference on Robot Learning, pages 1036–1049. PMLR, 2023
2023
-
[29]
Z. Liu, C. Chi, E. Cousineau, N. Kuppuswamy, B. Burchfiel, and S. Song. Maniwav: Learning robot manipulation from in-the-wild audio-visual data. In8th Annual Conference on Robot Learning, 2024
2024
-
[30]
R. Feng, J. Hu, W. Xia, T. Gao, A. Shen, Y . Sun, B. Fang, and D. Hu. Anytouch: Learning unified static-dynamic representation across multiple visuo-tactile sensors.arXiv preprint arXiv:2502.12191, 2025
2025 arXiv
-
[31]
Gong, Y .-A
Y . Gong, Y .-A. Chung, and J. Glass. Ast: Audio spectrogram transformer.arXiv preprint arXiv:2104.01778, 2021
2021 arXiv
-
[32]
Niizumi, D
D. Niizumi, D. Takeuchi, Y . Ohishi, N. Harada, and K. Kashino. Byol for audio: Exploring pre-trained general-purpose audio representations.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:137–151, 2023. doi:10.1109/TASLP.2022.3221007
2023
-
[33]
Morgado, N
P. Morgado, N. Vasconcelos, and I. Misra. Audio-visual instance discrimination with cross- modal agreement. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12475–12486, 2021
2021
-
[34]
H. Li, Y . Zhang, J. Zhu, S. Wang, M. A. Lee, H. Xu, E. Adelson, L. Fei-Fei, R. Gao, and J. Wu. See, hear, and feel: Smart sensory fusion for robotic manipulation. In K. Liu, D. Kulic, and J. Ichnowski, editors,Proceedings of The 6th Conference on Robot Learning, volume 205 of...
2023
-
[35]
Nagrani, S
A. Nagrani, S. Yang, A. Arnab, A. Jansen, C. Schmid, and C. Sun. Attention bottlenecks for multimodal fusion. In M. Ranzato, A. Beygelzimer, Y . Dauphin, P. Liang, and J. W. Vaughan, editors,Advances in Neural Information Processing Systems, volume 34, pages 14200–14213. Curra...
2021
-
[36]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale.arXiv preprint arXiv:2010.11929, 2020
2010 arXiv
-
[37]
C. Chi, Z. Xu, C. Pan, E. Cousineau, B. Burchfiel, S. Feng, R. Tedrake, and S. Song. Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots. InProceedings of Robotics: Science and Systems (RSS), 2024
2024
-
[38]
Young, D
S. Young, D. Gandhi, S. Tulsiani, A. Gupta, P. Abbeel, and L. Pinto. Visual imitation made easy. In J. Kober, F. Ramos, and C. Tomlin, editors,Proceedings of the 2020 Conference on Robot Learning, volume 155 ofProceedings of Machine Learning Research, pages 1992–2005. PMLR, 16...
2020
-
[39]
N. M. M. Shafiullah, A. Rai, H. Etukuru, Y . Liu, I. Misra, S. Chintala, and L. Pinto. On bringing robots home.arXiv preprint arXiv:2311.16098, 2023
2023 arXiv
-
[40]
Caron, H
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin. Emerging properties in self-supervised vision transformers. InProceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021
2021
-
[41]
X. Chen, M. Ding, X. Wang, Y . Xin, S. Mo, Y . Wang, S. Han, P. Luo, G. Zeng, and J. Wang. Context autoencoder for self-supervised representation learning.International Journal of Computer Vision, 132(1):208–223, 2024
2024
-
[42]
Huang, X
H.-J. Huang, X. Guo, and W. Yuan. Understanding Dynamic Tactile Sensing for Liquid Property Estimation. InProceedings of Robotics: Science and Systems, New York City, NY , USA, June
-
[43]
C. Matl, Y . Narang, R. Bajcsy, F. Ramos, and D. Fox. Inferring the material properties of granular media for robotic tasks. In2020 IEEE International Conference on Robotics and Automation (ICRA), pages 2770–2777, 2020. doi:10.1109/ICRA40945.2020.9197063
2020
-
[44]
Liang, S
H. Liang, S. Li, X. Ma, N. Hendrich, T. Gerkmann, F. Sun, and J. Zhang. Making sense of audio vibration for liquid height estimation in robotic pouring. In2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 5333–5339, 2019. doi:10.1109/ IROS4...
2019
-
[45]
Clarke, T
S. Clarke, T. Rhodes, C. G. Atkeson, and O. Kroemer. Learning audio feedback for estimating amount and flow of granular material. In A. Billard, A. Dragan, J. Peters, and J. Morimoto, editors,Proceedings of The 2nd Conference on Robot Learning, volume 87 ofProceedings of Machi...
2018
-
[46]
Sharma, C
A. Sharma, C. Higuera, C. K. Bodduluri, Z. Liu, T. Fan, T. Hellebrekers, M. Lambeta, B. Boots, M. Kaess, T. Wu, F. R. Hogan, and M. Mukadam. Self-supervised perception for tactile skin covered dexterous hands, 2025. URLhttps://arxiv.org/abs/2505.11420
2025 arXiv
-
[47]
B. Tang, M. A. Lin, I. Akinola, A. Handa, G. S. Sukhatme, F. Ramos, D. Fox, and Y . Narang. Industreal: Transferring contact-rich assembly tasks from simulation to reality.arXiv preprint arXiv:2305.17110, 2023
2023 arXiv
-
[48]
Y . Wu, Z. Chen, F. Wu, L. Chen, L. Zhang, Z. Bing, A. Swikir, S. Haddadin, and A. Knoll. Tacdiffusion: Force-domain diffusion policy for precise tactile manipulation.arXiv preprint arXiv:2409.11047, 2024
2024 arXiv
-
[49]
T. Z. Zhao, V . Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware.arXiv preprint arXiv:2304.13705, 2023
2023 arXiv
-
[50]
Kumar, Z
A. Kumar, Z. Fu, D. Pathak, and J. Malik. Rma: Rapid motor adaptation for legged robots. arXiv preprint arXiv:2107.04034, 2021
2021 arXiv
-
[51]
H. Qi, A. Kumar, R. Calandra, Y . Ma, and J. Malik. In-Hand Object Rotation via Rapid Motor Adaptation. InConference on Robot Learning (CoRL), 2022
2022
-
[52]
Zhang, A
L. Zhang, A. Rao, and M. Agrawala. Adding conditional control to text-to-image diffusion models. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 3836–3847, October 2023
2023
-
[53]
K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770– 778, 2016. 12
2016
-
[54]
Cadene, S
R. Cadene, S. Alibert, A. Soare, Q. Gallouedec, A. Zouitine, and T. Wolf. Lerobot: State-of-the- art machine learning for real-world robotics in pytorch. https://github.com/huggingface/ lerobot, 2024. 13 Appendix A Datasets A.1 Dataset forSparsh-XSSL Pretraining 272min270min 2...
2024
-
[56]
Tap-Foamwork
Golf-Circular-Fabric 1.Golf-Circular-Foamwork 2.Golf-Circular-Grass 3.Golf-Circular-Plastic 4.Golf-Slide-Fabric 5.Golf-Slide-Foamwork 6.Golf-Slide-Grass 7.Golf-Slide-Plastic 8.Golf-Tap-Fabric 9.Golf-Tap-Foamwork 10.Golf-Tap-Grass 11.Golf-Tap-Plastic 12.LEGO-Circular-Fabric 13....
-
[2022]
doi:10.15607/RSS.2022.XVIII.072
2022 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.