{"id":"b53ace31-eb63-46bf-8f2b-84750a510afe","arxiv_id":"2508.09836","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A modular robotic skin combined with an unsupervised deep state-space model shows that late-fusion multimodal tactile sensing and moderate palpation motions best infer soft object properties.","lead":"This paper tests how a robotic skin's stiffness, sensing channels, and poking motions affect a robot's ability to infer the softness, texture, and shape of soft objects. It introduces the latent filter, an unsupervised deep state-space model, and shows that combining force and vibration sensors beats using either alone.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Sequence-level train/test split lets the regressor memorize object identity, so MULTI-L's advantage and the 'generalizable/causal' claims are not yet established.","rationale":"The reader's weakest assumption correctly identifies the same load-bearing concern: the random sequence-level split permits object-specific leakage, undermining the generalizability and causality claims. I examined the ELBO derivation and the model architecture; no separate internal inconsistency emerged. The MULTI-L advantage may still be real, but the current evaluation cannot distinguish between memorizing object-specific latent patterns and genuinely inferring soft-object properties. This concern is central because the paper's stated contribution is a generalizable, causal representation for tactile perception, and the headline claim about late fusion depends on the same flawed protocol. A conditional verdict is appropriate: the findings are plausible but not yet established for new objects. The recommended concrete check is an object-disjoint split, which directly tests whether the reported NMSE rankings survive when test objects were never seen during training.","tokens_in":15326,"tokens_out":3061,"duration_ms":36985,"concrete_test":"Retrain the full pipeline (latent filter plus kernel-ridge regressor) using an object-disjoint split: hold out, e.g., 8 of the 32 objects, train only on interactions of the remaining 24 objects, and evaluate on the held-out objects. Report NMSE per property, modality, and interaction primitive across at least 3 random seeds. If MULTI-L's advantage over FSRtb/ACC does not survive object-disjoint evaluation, or if absolute NMSE increases sharply, the Figure 5 result reflects object-specific memorization rather than generalizable property inference.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section 2.6 ('MULTI-L consistently yields the lowest NMSE across all interaction primitives') is only as strong as the evaluation protocol. Section 4.2 states that 25% of the dataset was reserved for testing, but the split is at the level of interaction sequences, not objects. Since the dataset contains only 32 objects, each probed repeatedly under 16 action settings, the same physical objects appear in both training and test folds. The kernel-ridge regressor used in Section 2.6 can therefore latch onto object-specific latent patterns (effectively object identity) rather than the intended generative factors (spatial frequency, amplitude, stiffness). This is compounded by the hierarchical prior p(yt|at, N) in Section 2.3, which conditions on the soft-object label N; object identity is thus explicitly available during latent-filter training. Calling the model unsupervised is imprecise, and the reported NMSE partly measures object memorization rather than property-level generalization. With this protocol, the abstract's claims of 'generalizable and causal inference' and the conclusion that late fusion is the best sensing strategy for estimating soft-object properties are not supported. No object-disjoint validation or statistical significance testing is reported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies how sensor embodiment, sensing modality, and interaction strategy jointly shape tactile perception in robots. It introduces a modular e-Skin with two stiffness variants (soft Ecoflex, hard DragonSkin) and three sensing channels (accelerometers, two FSR layers), a curated set of 32 wave objects with controlled amplitude, spatial frequency, stiffness, and heterogeneity, and three palpation primitives (pressing, precession, sliding). The central modeling contribution is the \"Latent Filter,\" an action-conditioned deep state-space model whose latent state is partitioned into directly observable (z_t) and indirectly observable (y_t) parts, trained with a variational ELBO. The learned latent representations are evaluated with kernel ridge regression to predict ground-truth object properties, reporting NMSE over time for different modality configurations, interaction parameters, and skin stiffnesses. The main claims are that late-fusion multimodal encoding (MULTI-L) consistently yields the lowest NMSE, that moderately high interaction frequencies and larger indentation depths improve inference, and that soft versus hard skin offers complementary advantages depending on object softness and interaction type.","tokens_in":15598,"tokens_out":3376,"duration_ms":40327,"significance":"If the empirical claims survive a more rigorous evaluation, the paper would make a useful contribution: it provides a systematic, hardware-grounded study of how skin compliance, multimodal sensing, and action choice affect tactile perception, and it proposes a latent-variable modeling framework that goes beyond static feature extraction. The curated object set and the release of code and a representative data subset are strengths. The time-resolved NMSE analysis is a constructive way to compare sensing strategies. However, the current evaluation protocol does not establish the central generalization and causality claims: the train/test split is not object-disjoint, the hierarchical prior uses object labels during latent-filter training, and the regression evaluation appears to train and test on the same interaction sequences. The claimed superiority of MULTI-L and the interaction-parameter conclusions rest on this protocol, so they are not yet convincing.","major_comments":[{"comment":"The evaluation protocol is the main weakness. Section 4.2 states only that \"Twenty-five percent of the dataset was reserved for testing,\" but the split is at the level of interaction sequences, not objects. With only 32 objects and 16 action settings per object, the same physical objects appear in both training and test folds. The kernel-ridge regressor in Section 2.6 can therefore latch onto object-specific latent patterns (effectively object identity) rather than the intended generative factors (spatial frequency, amplitude, stiffness). This undermines the abstract's claims of \"generalizable and causal inference\" and the Section 2.6 conclusion that MULTI-L is the best sensing strategy. Moreover, the regression methodology described in Section 2.6 -- training on the final segments of each interaction sequence and applying to samples from the evolving latent space -- suggests that test s","section":"Section 2.6 / Section 4.2"},{"comment":"The model is called \"unsupervised\" in the abstract, introduction, and Section 2.3, but the learnable hierarchical prior p(y_t | a_t, N) explicitly conditions on N, described as \"the soft-object label.\" This means object identity is available during latent-filter training. The model may therefore be label-conditioned or semi-supervised rather than unsupervised. This is not merely a terminological issue: the label-conditioned prior could allow the latent variable y_t to encode object identity directly, which is directly relevant to the leakage concern above. The claim should be revised, and an ablation without the N-conditioned prior should be reported to support the \"unsupervised\" and \"causal\" interpretation.","section":"Section 2.3 / Abstract"},{"comment":"The quantitative comparisons lack statistical support. Figure 5 shows NMSE curves with standard deviation bands, but the caption states that the standard deviation is scaled by a factor of 10 \"for visual clarity\" without justification, and no number of seeds or repeated runs is reported. Tables 1 and 2 report single NMSE values per action parameter and per object group with no confidence intervals or significance tests. Consequently, the claimed consistent advantage of MULTI-L over FSRtb, and the differences between interaction parameters, could be within run-to-run variability. I request repeated training runs (or at least bootstrap confidence intervals over sequences/objects) and significance testing for the central comparisons.","section":"Table 2 / Section 2.8"},{"comment":"Table 2 shows that for heterogeneous objects, the NMSE for heterogeneity is very high (e.g., 0.516, 0.376, 0.629 for soft skin; 0.543, 0.685, 1.172 for hard skin), yet the text emphasizes that soft skin is \"particularly well-suited for perceiving the properties of heterogeneous materials.\" An NMSE above 0.5 for the very property that defines the heterogeneous class suggests the model fails to estimate heterogeneity reliably. This tension should be discussed explicitly; the current narrative overstates the positive result and underplays a clear failure case.","section":"Table 2 / Section 2.8"}],"minor_comments":[{"comment":"The ELBO expressions contain unmatched brackets and the KL terms are written without closing parentheses: e.g., \"KL[qf ilt(zt)||p(zt|zt−1, yt, at)\" is missing a closing parenthesis. This makes the derivation harder to follow.","section":"Eq. (3) and Eq. (13)"},{"comment":"Typo: \"ACC can detect features by sliding that FSRs alone cannot not distinguish\" should be \"cannot distinguish.\"","section":"Section 2.5"},{"comment":"The sentence \"The tactile perception through sensors embodiment and the latent filter\" is incomplete and should be rewritten.","section":"Section 3"},{"comment":"The scaling of standard deviations by a factor of 10 is not explained. If the bands are small, report them unscaled or use a log scale; if they are scaled for visibility, state why and indicate the true magnitude.","section":"Figure 5 caption"},{"comment":"The preprocessing timing is inconsistent: accelerometer data is said to be downsampled to 600 Hz yielding 6000 samples, while FSR data is downsampled to 30 Hz yielding 300 samples. The ratio is 20x, which is consistent, but the text should clarify that the 6000 and 300 refer to the same 10-second window. Also, \"hop length of 20\" in the spectrogram is unusual given a window size of 800; please define units.","section":"Section 4.2"},{"comment":"Some references are incompletely formatted (e.g., missing publisher location in [4]) and the GitHub repository URL is not checked. Please ensure all references are complete and the repository is accessible.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The central empirical claims rest on an evaluation protocol that does not currently support them. The fix -- object-disjoint validation and clarification of the regression split -- is within the scope of a revision, so I do not recommend rejection. I would also ask the editor to verify the novelty boundary with respect to the authors' own prior work on the e-Skin (reference [8], accepted at WHC 2025), since the present paper's sensor contribution may substantially overlap with that publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this one. First, the empirical core is genuinely useful: a modular e-Skin with tunable compliance, a new wave-object dataset with controlled surface and bulk properties, and a systematic comparison of modalities, skin stiffness, and three palpation primitives. The finding that late fusion of acceleration and force channels beats early fusion and single modalities is plausible and consistent with the different statistics of those signals. Second, the headline claims — \"generalizable and causal inference\" — are not supported by the evaluation as run. The train/test split is by interaction sequence, not by object, and the 32 objects all appear in both folds. The kernel-ridge regressor in Section 2.6 can therefore latch onto object identity rather than the intended physical properties.\n\nThe latent filter itself is a reasonable extension of Deep Variational Bayes Filters and hierarchical deep state-space models, with the z/y observability partition being the main novelty. Calling it unsupervised is imprecise: Section 2.3 conditions the prior on the soft-object label N, so object identity is available during latent training. That is not fatal — the NMSE numbers are genuine predictions from held-out sequences — but it does undercut the \"causal\" language and the abstract's generalization claim.\n\nWhat I credit: the ablation coverage is unusually thorough for this area. Six modality configurations, two skin stiffnesses, sixteen interaction parameters per primitive, and Table 2's object-group breakdown tell a coherent story. The paper also flags its own limitations in places, e.g. the Euclidean-distance caveat in Section 2.5. Code is public; the full dataset is not, and the e-Skin hardware sits in an unpublished companion paper, which caps reproducibility for now.\n\nThe soft spots, in proportion: (1) no object-disjoint validation, the main issue; (2) no significance tests or error bars on the tables, and Figure 5's ×10 SD scaling is unexplained — minor but easy to fix; (3) the causal claim is rhetorical, not demonstrated; (4) the paper does not say explicitly whether the latent filter was trained only on the 75% split or on all data — if the latter, there is additional leakage into the latent space, and that should be clarified.\n\nWho this is for: robotics and tactile-sensing researchers deciding how to co-design skin compliance, sensing channels, and palpation policies. It deserves a serious referee; an editor should send this out rather than desk-reject. My recommendation: engage with it, ask for an object-level split or at least per-object cross-validation, multi-seed runs with significance tests, and softened causal language. The empirical backbone would survive that revision, and the paper would be stronger for it.","headline":"A genuinely useful systematic ablation of how skin stiffness, sensing modality, and palpation strategy shape tactile soft-object estimation, wrapped in a latent-filter model whose generalizable/causal claims outrun the sequence-level evaluation protocol.","tokens_in":16092,"tokens_out":3997,"would_cite":true,"duration_ms":39673,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that a modular e-Skin that encodes force and vibration in separate pathways and fuses them late, within an action-conditioned latent state-space model, estimates soft object properties more accurately than any single sensi","keywords":["tactile perception","e-Skin","multimodal sensing","deep state-space model","latent filter","palpation","soft robotics","viscoelasticity"],"falsifier":"Retrain the latent filter and the kernel ridge regressor on a random subset of objects (e.g., all sequences from objects 0-13) and test on the remaining objects (14-31, including the heterogeneous set), reporting NMSE for MULTI-L against single modalities. If the late-fusion advantage disappears or errors rise sharply compared with the reported sequence-split numbers, the generalization and modality-superiority claims are artefacts of object leakage. A second decisive check is to shuffle object labels while keeping the same tactile sequences: if the latent space still separates objects rather","tokens_in":15196,"feed_emoji":"🖐️","tokens_out":9493,"duration_ms":100985,"temperature":0.7,"pith_summary":"This paper sets out to show that tactile perception of soft objects is a joint product of the sensing skin's mechanics, the sensor channels it carries, and the way the robot moves, and that all three can be unified in one learning model. The authors build a modular e-Skin whose silicone stiffness can be switched between soft and hard and whose readings include normal forces, shear-proxy forces, and vibrations, then palpate a set of wave objects with controlled stiffness, surface texture, and heterogeneity using pressing, precession, and sliding. They propose the Latent Filter, an unsupervised, action-conditioned deep state-space model that folds temporal interaction dynamics into a structured latent space, and show that a late-fusion combination of force and vibration channels yields the lowest normalized mean squared error across all interaction primitives. The result matters because robots that manipulate deformable objects—in medical palpation, food handling, or delicate assembly—need to infer physical properties from active touch rather than from pre-labelled categories.","feed_headline":"Fuse force and vibration late to read soft objects best","feed_subtitle":"Separate encoders for force and vibration channels beat single sensors and early fusion in every palpation primitive.","key_machinery":"The Latent Filter is an unsupervised, action-conditioned deep state-space model that factorizes the latent state into a directly observable component $z_t$ and an indirectly observable component $y_t$, with an LSTM approximating the latter, Bayesian integration of an inverse measurement model with a recursive dynamics model, and a learnable hierarchical prior conditioned on action and object label that pushes the latents toward causal mechanical properties. Its companion is the modular e-Skin: an accelerometer array for high-frequency vibration and two FSR arrays separated by a compliant interlayer for normal and differential (shear-proxy) forces, in soft (Ecoflex) or hard (DragonSkin) silic","core_discovery":"The paper's central result is that the late-fusion multimodal configuration (MULTI-L), in which the accelerometer layer and the two force-sensing-resistor layers are encoded by separate networks and combined only at the latent level, consistently achieves the lowest normalized mean squared error for estimating the wave objects' spatial frequency, amplitude, stiffness, and heterogeneity, across pressing, precession, and sliding. Two further findings carry the argument: the two FSR layers are the main contributors for stiffness and heterogeneity because differential normal forces proxy shear and skin stretch, while the accelerometer contributes transient information that helps most in sliding;","pith_inferences":["A directly testable extension is an object-split evaluation: train the latent filter on a subset of the objects and test on unseen objects. Because the current 75/25 split separates interaction sequences rather than objects, such an evaluation would show whether the late-fusion advantage and the causal-claim survive truly novel objects.","The causal reading of the latent dimensions depends on the supervised alignment used for evaluation; a natural extension is to probe latent traversals directly and check whether they monotonically track each generative parameter without object identity.","The specific advantage of soft skin on heterogeneous objects suggests an adaptive-compliance skin that stiffens or softens in real time could outperform either fixed configuration across the full object set.","A multi-skin comparison with varied silicone thicknesses or artificial fingerprints could separate the contribution of mechanical filtering from the sensor channel itself, which the current two-point stiffness comparison cannot resolve."],"forward_implications":["Late fusion of force and vibration channels, rather than early concatenation, should be the default architecture for multimodal tactile encoders in soft-robot perception.","Palpation trajectories should be chosen by target property: moderate frequencies around 0.6 Hz with sufficient indentation depth improve estimation, while pressing is relatively insensitive to parameter choice beyond convergence speed.","Skin stiffness should be co-designed with task: softer skins for heterogeneous or compliant objects and sliding interactions, harder skins for bulk stiffness and surface geometry.","Unsupervised, action-conditioned latent dynamics can replace handcrafted features and fixed category classifiers for soft-object property regression, enabling continuous estimation of multiple physical properties at once."],"supporting_citations":[{"why":"supplies the modular, multi-modal e-Skin hardware used in every experiment.","marker":"[8]"},{"why":"established that a soft outer layer acts as a mechanical low-pass filter on tactile signals, motivating the embodiment hypothesis tested here.","marker":"[5,6]"},{"why":"provides the deep state-space modelling and learnable hierarchical prior on which the latent filter is built.","marker":"[33]"},{"why":"supplies the deep variational Bayes filtering formalism, the annealing strategy, and the supervised latent-alignment regression used for evaluation.","marker":"[35]"},{"why":"shows that multi-step dynamic interactions can reveal dense physical object representations, motivating the action-conditioned temporal modeling.","marker":"[19]"},{"why":"prior work on action augmentation for soft-body palpation that this paper's interaction primitives and dataset extend.","marker":"[10]"},{"why":"the human multisensory fusion result that motivates expecting multimodal tactile integration to outperform single channels.","marker":"[37]"},{"why":"the account of haptic exploratory procedures that motivates pressing, precession, and sliding as task-relevant palpation primitives.","marker":"[38]"}],"fun_headline_variants":["Late fuse force and vibration for best soft-object reading","Separate encoders beat early fusion in tactile palpation","Soft object perception: late multimodal fusion wins","Force layers read stiffness; accelerometer boosts sliding","Multimodal tactile sensing: keep force and vibration apart"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that evaluating the model on held-out interaction sequences of the same physical objects measures generalization; if instead the objects themselves must be held out, the model may be memorizing object-specific latent patterns rather than learning the mechanical properties it claims to infer.","fun_headline_variants_meta":{"raw":{"variants":["Late fuse force and vibration for best soft-object reading","Separate encoders beat early fusion in tactile palpation","Soft object perception: late multimodal fusion wins","Force layers read stiffness; accelerometer boosts sliding","Multimodal tactile sensing: keep force and vibration apart"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000718,"raw_usage":{"total_tokens":3035,"prompt_tokens":692,"completion_tokens":2343,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":436,"completion_tokens_details":{"reasoning_tokens":2268}},"tokens_in":436,"tokens_out":2343,"duration_ms":19136,"temperature":1.0,"reasoning_tokens":2268,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:46:36.605149+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the latent filter and the kernel ridge regressor on a random subset of objects (e.g., all sequences from objects 0-13) and test on the remaining objects (14-31, including the heterogeneous set), reporting NMSE for MULTI-L against single modalities. If the late-fusion advantage disappears or errors rise sharply compared with the reported sequence-split numbers, the generalization and modality-superiority claims are artefacts of object leakage. A second decisive check is to shuffle object labels while keeping the same tactile sequences: if the latent space still separates objects rather","supporting_citations":[{"cited_title":"In: IEEE World Haptics Conference (WHC) (2025)","cited_arxiv_id":null,"evidence_quote":"supplies the modular, multi-modal e-Skin hardware used in every experiment."},{"cited_title":"In: Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P.S., Vaughan, J.W","cited_arxiv_id":null,"evidence_quote":"provides the deep state-space modelling and learnable hierarchical prior on which the latent filter is built."},{"cited_title":"In: Proceedings of International Conference on Learning Representations (ICLR) (2017)","cited_arxiv_id":null,"evidence_quote":"supplies the deep variational Bayes filtering formalism, the annealing strategy, and the supervised latent-alignment regression used for evaluation."},{"cited_title":"In: Proc","cited_arxiv_id":null,"evidence_quote":"shows that multi-step dynamic interactions can reveal dense physical object representations, motivating the action-conditioned temporal modeling."},{"cited_title":"Soft Robotics 9(2), 280–292 (2022) https://doi.org/10.1089/soro.2020.0129","cited_arxiv_id":null,"evidence_quote":"prior work on action augmentation for soft-body palpation that this paper's interaction primitives and dataset extend."},{"cited_title":"Science 298(5598), 1627–1630 (2002)","cited_arxiv_id":null,"evidence_quote":"the human multisensory fusion result that motivates expecting multimodal tactile integration to outperform single channels."},{"cited_title":"Acta psychologica 84(1), 29–40 (1993)","cited_arxiv_id":null,"evidence_quote":"the account of haptic exploratory procedures that motivates pressing, precession, and sliding as task-relevant palpation primitives."}],"review_version":1}