{"id":"8676187a-2ae4-4e0a-8bd3-5ad6c949687c","arxiv_id":"2607.11018","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Tuple-and-spline reductions plus lightweight rectified flow ground MLLM responses into soft-trunk motions, raising held-out correctness to 77.2% and overall satisfaction to 82%.","lead":"A dual-reduction plus rectified-flow pipeline turns open-vocabulary MLLM responses into whole-body tendon trajectories for an elephant-inspired soft trunk. Ablations and a 100-person HRI study show large gains in held-out motion match and user satisfaction versus simple baselines.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Held-out RMS success is circular: the same MLLM-screened library that trains the model also defines the gold targets and the τ_rms=0.35 threshold.","rationale":"The reader correctly isolates the weakest link: the 256-trajectory library and the RMS≤0.35 rule jointly define both training targets and the success metric. My concern is the same circularity made fully explicit—the held-out set is not an external gold standard but a held-out slice of the same MLLM-screened manifold. The engineering pipeline (dual reduction + rectified flow) is cleanly ablated and the hardware/user-study results are real, so the contribution remains credible inside its niche; the evaluation simply does not yet support the stronger claim of generalizable “grounding correctness.” No change of verdict is required; the existing CONDITIONAL already captures the gap. A blinded independent human rating of physical executions would settle whether the numbers survive outside the library.","tokens_in":12274,"tokens_out":618,"duration_ms":6437,"concrete_test":"Collect an independent test set of ≥30 new open-vocabulary responses never seen by the library construction process; have 3–5 naïve human raters (not the original annotators) score physical executions of (a) library references, (b) RF samples, and (c) DDPM samples for social appropriateness and physical feasibility on a 5-point scale, blinded to source. If RF’s human-positive rate falls below ~70% or loses its edge over DDPM, the 77.2% claim does not transfer beyond the library.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central quantitative claim (Table II: 25.0%→77.2% held-out correctness) rests on treating the 64 held-out library trajectories as independent gold standards for “socially valid whole-body motion.” Those trajectories were produced by the identical pipeline used at train time (MLLM reduction under C, Catmull–Rom fitting, fixed-prompt MLLM screening for semantic agreement, and manual verification; Sec. III-C.1). The success threshold τ_rms=0.35 was itself “calibrated against human annotations and LLM-as-a-judge assessments of the corresponding hardware motions” (Sec. IV-A.3). Consequently the metric largely measures how well the generator reproduces the library manifold rather than how well it produces independently judged feasible, socially appropriate trunk motions. If the library is narrow or the threshold is library-tuned, both the ablation percentages and the claim of feasible grounding become self-referential. The HRI study only shows that any trunk motion beats AV-only; it does not break this circularity.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes a whole-body semantic-to-actuation grounding pipeline for elephant-inspired soft-trunk HRI. Open-vocabulary MLLM responses are reduced to morphology-aligned intent–intensity tuples (16 intents × 4 intensity levels), continuous tendon trajectories are parameterized by compact Catmull–Rom control matrices (Nu=4, P=4), and a lightweight rectified-flow model samples from the conditional distribution over those controls. Ablations on a 256-sample library (Table II) show progressive gains: Raw+Dense+CNN 25.0% held-out success → Tuple+Spline+RF 77.2%, with RF also faster than DDPM (4.87 ms vs 7.86 ms) while retaining diversity. A 100-participant within-subject physical study reports that adding the generated trunk channel raises positive overall-satisfaction ratings from 46% to 82% versus audiovisual-only interaction.","tokens_in":12629,"tokens_out":1071,"duration_ms":8670,"significance":"If the claims hold, the work supplies a practical dual-reduction interface that bridges open-vocabulary MLLM responses and high-dimensional continuum actuation for close-contact social HRI—an under-served setting relative to tip-centric VLA. The clean ablation isolation of semantic reduction, spline reduction, and RF versus CNN/DDPM, the explicit one-to-many sampling, and the sizable physical user study are concrete strengths. The contribution is primarily systems-level rather than a new theoretical result, but it is timely for soft-robot HRI and demonstrates deployable inference latency.","major_comments":[{"comment":"Sec. III-C.1 and IV-A.3 / Table II: Held-out correctness (normalized RMS ≤ τ_rms=0.35) is measured against library trajectories that were themselves produced by the same MLLM reduction under C, Catmull–Rom fitting, fixed-prompt MLLM screening, and manual verification used at train time. The threshold was calibrated on the same hardware motions. This makes the 25.0%→77.2% claim largely a measure of library-manifold reproduction rather than independently judged social validity. The progressive ablations and DDPM comparison remain informative, but the central quantitative claim needs either an external human/LLM-as-judge evaluation of generated motions or an explicit statement that success is library-relative.","section":"Sec. III-C.1, IV-A.3, Table II"},{"comment":"Sec. IV-B: The physical HRI study compares AV-only versus AV+generated trunk. It therefore shows that adding any trunk motion improves satisfaction (46%→82%), not that the proposed grounding is preferable to scripted or alternative generators. The manuscript itself notes that a direct perceptual comparison with manually scripted motions is left for future work; without that (or at least a no-grounding trunk baseline), the user-study claim cannot be read as validation of the semantic-to-actuation pipeline specifically.","section":"Sec. IV-B"}],"minor_comments":[{"comment":"Eq. (3) is presented as a conceptual argmin but is never optimized; the practical implementation is the few-shot MLLM call in Eq. (4). Clarify that the former is only notational.","section":"Sec. III-A.2"},{"comment":"Fig. 4 qualitative labels (Wrong/Marginal/Correct) are useful but the corresponding RMS values are given only for selected cases; reporting RMS for every panel would strengthen the visual–quantitative link.","section":"Fig. 4"},{"comment":"Diversity (Eqs. 15–16) is mean pairwise RMS; a brief note on whether higher diversity is always desirable (vs. occasional outliers noted for DDPM) would help interpretation of the 0.148 vs 0.165 comparison.","section":"Sec. IV-A.3"},{"comment":"Table I lists free parameters (P=4, S=50, τ_rms=0.35, network width) without sensitivity analysis; a short appendix or sentence on robustness to P and S would be useful.","section":"Table I"},{"comment":"Minor wording: Abstract and Introduction use both “intent-intensity” and “intent–intensity”; standardize the en-dash.","section":"Abstract / Introduction"}],"recommendation":"major_revision","confidential_remarks":"The circularity concern raised by the stress-test is real and load-bearing for the ablation percentages, but it is fixable by adding an independent perceptual evaluation of generated motions (or by reframing success as library-relative). The paper is a solid systems contribution for soft-robot HRI; I would not reject on novelty grounds. Scope fits cs.RO / soft robotics venues well."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a solid engineering pipeline paper, not a theory paper. They take open-vocab MLLM responses, force them into a small morphology-aligned intent–intensity space (16×4), compress tendon trajectories to Catmull–Rom controls, and sample with lightweight rectified flow. That dual reduction for whole-body soft-trunk HRI is the actual new piece; RF, Catmull–Rom, and few-shot MLLM reduction are each known and cited.\n\nWhat works: Table II is clean. You can watch each stage move the needle—raw dense CNN 25%, tuple 48.8%, spline 56.2%, DDPM 71.9%, RF 77.2%—with RF also faster (4.87 ms vs 7.86 ms) and diversity still present. Qualitative tendon plots match the story. The 100-participant within-subject study is large for HRI and shows a clear user-facing lift when trunk motion is added (overall satisfaction 46%→82%). Hardware is real; the problem framing (tip VLA underspecifies body shape; dense whole-body is hard) is right.\n\nSoft spots, in proportion: the 256-sample library is small, and held-out success is RMS to library trajectories built with the same reduction, MLLM screening, and manual checks. τ_rms=0.35 was calibrated on those motions. So the metric partly measures reproduction of their own manifold. That is mild circularity, not a collapse—the progressive ablations and DDPM comparison still hold inside the same setup, and they checked ordering across several thresholds. The HRI arm only shows “any trunk motion beats AV-only,” not that this generator beats scripted or alternative generators. No code or data. Free parameters (P, S, τ_rms, library size) are ordinary for this class of work.\n\nWho it’s for: soft robotics and close-contact social HRI people who need an executable whole-body channel. Outside that niche it is optional. Math is standard flow matching; citations look appropriate; no internal contradiction.\n\nSend it to peer review. Gaps are fixable with a larger independent library, a human/LLM judge not tied to the training set, and a generator-vs-generator user arm. Worth engaging if you work on soft trunks or embodied social response; not a must-read otherwise.","headline":"Usable dual-reduction + rectified-flow stack for whole-body soft-trunk social motion; ablations and 100-person study are real, but held-out “correctness” is partly library-circular.","tokens_in":13229,"tokens_out":605,"would_cite":false,"duration_ms":16631,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Soft-trunk robots can turn open language into whole-body motion by reducing intent and actuation, then sampling with rectified flow.","keywords":["soft robotics","whole-body semantic-to-actuation grounding","bio-inspiration","MLLM","HRI","rectified flow","Catmull-Rom spline","continuum manipulators"],"falsifier":"If independent judges or a larger held-out library show that many high-RMS samples are still perceived as valid trunk gestures, or that many low-RMS samples look wrong on the real platform, the reported correctness gains and the social-validity claim would fail.","tokens_in":13203,"feed_emoji":"🐘","tokens_out":677,"duration_ms":6262,"temperature":0.7,"pith_summary":"Close-contact human-robot interaction needs bodies that can express intent through shape, not only through speech or a screen. Soft elephant-inspired trunks can do that, but ordinary language-to-action pipelines fail them: tip commands do not specify whole-body posture, and raw high-dimensional tendon trajectories are hard to keep feasible. This paper claims a dual reduction plus lightweight flow matching solves the mismatch. Open multimodal responses are first mapped into a small set of morphology-aligned intent-intensity pairs; continuous tendon profiles are then encoded as compact Catmull-Rom spline controls; a rectified-flow model learns the conditional distribution over those controls and samples smooth, executable motions. Ablations show held-out grounding correctness rising from 25 percent for raw dense regression to 77.2 percent for the full pipeline, with faster inference than a diffusion baseline. In a 100-person study, adding the generated trunk channel lifted positive overall satisfaction from 46 percent to 82 percent versus audiovisual-only interaction.","feed_headline":"Soft trunks turn open language into whole-body motion","feed_subtitle":"Dual reduction plus rectified flow lifts grounding correctness to 77% and satisfaction to 82%","key_machinery":"Dual reduction plus rectified-flow grounding: morphology-aligned intent-intensity tuples (sixteen intents by four intensities) condition a lightweight velocity field over Catmull-Rom spline control matrices, so continuous tendon trajectories are reconstructed from low-dimensional samples rather than predicted densely.","core_discovery":"The authors establish that open-vocabulary multimodal responses can be grounded into feasible whole-body soft-trunk actuation by first canonicalizing language into bounded intent-intensity tuples, then representing trajectories with compact Catmull-Rom controls, and finally sampling from a tuple-conditioned rectified-flow model. That pipeline yields higher held-out correctness than raw-response dense regression or a diffusion generator, faster inference than diffusion, retained diversity, and a clear user-facing gain when the trunk channel is added to audiovisual interaction.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Flow matching grounds language into soft-trunk whole-body motion","Soft trunks map open vocabulary to feasible body actuation","Lightweight flow turns MLLM replies into trunk trajectories","Intent tuples plus rectified flow enable soft-trunk HRI","Elephant soft trunks ground semantics into whole-body motion"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The claim rests on treating a small offline library of 256 expert trajectories, filtered by the same reduction rules and a fixed RMS tolerance of 0.35, as the gold standard for what counts as correct and socially valid whole-body motion.","fun_headline_variants_meta":{"raw":{"variants":["Flow matching grounds language into soft-trunk whole-body motion","Soft trunks map open vocabulary to feasible body actuation","Lightweight flow turns MLLM replies into trunk trajectories","Intent tuples plus rectified flow enable soft-trunk HRI","Elephant soft trunks ground semantics into whole-body motion"]},"model":"grok-4.5","effort":"low","cost_usd":0.003202,"raw_usage":{"total_tokens":1151,"prompt_tokens":830,"num_sources_used":0,"completion_tokens":61,"cost_in_usd_ticks":32020000,"prompt_tokens_details":{"text_tokens":830,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":260,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":830,"tokens_out":61,"duration_ms":2673,"temperature":1.0,"reasoning_tokens":260,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T07:36:39.232465+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"If independent judges or a larger held-out library show that many high-RMS samples are still perceived as valid trunk gestures, or that many low-RMS samples look wrong on the real platform, the reported correctness gains and the social-validity claim would fail.","supporting_citations":[],"review_version":1}