{"id":"c94722a4-247f-4c56-9321-adc4dbb8c00e","arxiv_id":"2506.12606","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Mamba-based HuBERT models match or exceed Transformer versions on speech tasks while using far less compute for long sequences and streaming ASR.","lead":"This paper tests Mamba state-space models as replacements for Transformers inside HuBERT-style self-supervised speech systems. If the efficiency and streaming gains hold, it offers a practical route to lower-cost long-audio and real-time speech processing.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Direct Mamba-for-Transformer swap in HuBERT pre-training may not preserve representation quality without speech-specific retuning of state size or discretization.","rationale":"The reader's weakest assumption directly identifies the same substitution risk. With the full manuscript now available, the concern remains load-bearing because the reported gains could depend on unstated adaptations rather than the Selective SSM itself. This moves the verdict from UNVERDICTED to CONDITIONAL pending the proposed check.","tokens_in":1652,"tokens_out":335,"duration_ms":45100,"concrete_test":"Re-run the HuBERT pre-training stage with the Mamba block using the exact hyper-parameters reported for the Transformer baseline (state dimension 16, no extra discretization tuning) and measure the final SUPERB ASR and speaker verification scores; if the Mamba variant falls more than 5% relative behind the Transformer on any causal task, the direct-substitution assumption does not hold.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim rests on Mamba-based HuBERT models delivering competitive SUPERB scores, higher-quality quantized units, and superior streaming ASR after straightforward substitution into the standard HuBERT pipeline. Speech inputs are continuous and high-rate; the selective SSM's discretization and state dimension (typically small) may fail to capture fine-grained acoustic structure unless the block is adapted. The paper reports better speaker-feature separation and causal performance, but this could be an artifact of the particular hyper-parameters chosen rather than an intrinsic property of the architecture. If the substitution works only after hidden retuning, the linear-time advantage for long-context fine-tuning does not generalize as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper explores Mamba-based alternatives to Transformer-based HuBERT models for speech self-supervised learning. It claims that these models achieve competitive performance on SUPERB probing benchmarks, especially in causal settings, produce higher-quality quantized representations, capture speaker-related features more distinctly, and enable efficient fine-tuning for long-context and streaming ASR due to the linear-time complexity of selective state space models.","tokens_in":1777,"tokens_out":464,"duration_ms":30017,"significance":"If the results are confirmed, this work is significant as it introduces an efficient architecture for speech SSL that scales better to long sequences and supports real-time applications. The open-source codebase at the provided GitHub link is a strength that facilitates reproducibility and further research in the community.","major_comments":[{"comment":"The description of the Mamba integration into the HuBERT pre-training pipeline does not include ablations on key hyperparameters such as state size or discretization step; without these, it is unclear whether the reported improvements require speech-specific retuning or hold under direct substitution as assumed.","section":"Section 3"},{"comment":"The SUPERB benchmark results are presented without reporting the number of independent runs, standard deviations, or statistical significance tests; this weakens the claim of competitive or superior performance in causal settings.","section":"Section 4.2"},{"comment":"The analysis of quantized units and speaker feature separation relies on qualitative observations or specific metrics; quantitative comparisons with baselines should be expanded to confirm higher quality.","section":"Section 5"}],"minor_comments":[{"comment":"The abstract mentions 'significantly lower compute' but does not quantify the savings; a brief mention of FLOPs or training time reduction would strengthen the claim.","section":"Abstract"},{"comment":"Ensure that all axes are clearly labeled and legends are legible for the streaming ASR performance plots.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The paper fits well within the scope of speech processing and self-supervised learning venues. The citation pattern appears standard, with appropriate references to Mamba and HuBERT works."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their detailed and constructive review of our manuscript. We address each of the major comments below and have incorporated revisions to improve the paper's clarity and completeness.","responses":[{"response":"We appreciate this comment. In designing our experiments, we deliberately used the default Mamba hyperparameters (state size d_state=16 and the standard discretization parameters from the Mamba paper) to test the hypothesis that Mamba can serve as a direct substitute for Transformers in speech SSL without requiring extensive speech-specific tuning. This approach aligns with our goal of exploring the architecture's potential in a straightforward manner. We have revised Section 3 to include explicit mention of these hyperparameter choices and a justification for not performing additional ablations in this initial exploration. We agree that further ablations would be beneficial and have added this as a suggested direction for future research.","revision_made":"partial","referee_comment":"[Section 3] The description of the Mamba integration into the HuBERT pre-training pipeline does not include ablations on key hyperparameters such as state size or discretization step; without these, it is unclear whether the reported improvements require speech-specific retuning or hold under direct substitution as assumed."},{"response":"We acknowledge the validity of this point. Given the high computational cost associated with pre-training large SSL models on speech data, our reported results are from single runs per configuration. We have updated Section 4.2 to clearly state that each result is from a single independent run and have included this information in the tables. While we did not conduct multiple runs or formal statistical tests, the trends observed across the various SUPERB tasks in causal settings are consistent and support our conclusions. We have moderated the language in the manuscript to reflect this and noted the limitation in the discussion section.","revision_made":"yes","referee_comment":"[Section 4.2] The SUPERB benchmark results are presented without reporting the number of independent runs, standard deviations, or statistical significance tests; this weakens the claim of competitive or superior performance in causal settings."},{"response":"Thank you for this suggestion. We have expanded Section 5 with additional quantitative analyses, including comparisons of codebook utilization rates and speaker identification accuracy using the quantized representations from both Mamba and Transformer models. These new metrics provide stronger quantitative evidence for the higher quality of the Mamba-based quantized units and better separation of speaker features. The revised section now includes direct numerical comparisons to the baseline.","revision_made":"yes","referee_comment":"[Section 5] The analysis of quantized units and speaker feature separation relies on qualitative observations or specific metrics; quantitative comparisons with baselines should be expanded to confirm higher quality."}],"tokens_in":1253,"tokens_out":582,"duration_ms":43875,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that this paper takes the standard HuBERT pre-training recipe and replaces the Transformer layers with Mamba blocks, then evaluates the resulting models on SUPERB probing tasks plus ASR fine-tuning. They report competitive scores overall, with particular strength in causal settings, better speaker-feature separation in the learned units, and lower compute when fine-tuning on long audio. The streaming ASR results are presented as an advantage too. The codebase release is useful for anyone who wants to check the implementation details.","headline":"Mamba can be swapped into a HuBERT pipeline for speech SSL and deliver efficiency wins on long-context and streaming ASR, but the direct-substitution story needs tighter controls on whether retuning was required.","tokens_in":2284,"tokens_out":184,"would_cite":false,"duration_ms":26562,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/ArithmeticFromLogic.lean","rs_theorem":"LogicNat induction and embed_strictMono_of_one_lt","paper_passage":"Mamba is a state space model (SSMs) whose discrete-time formulas are expressed as ht = Aht−1 + Bxt, yt = Cht"}],"headline":"Mamba SSM substitution in HuBERT pipeline has no overlap with RS cost/periodicity forcing","alignment":"orthogonal","rationale":"Paper's core machinery is empirical replacement of Transformer blocks by selective SSMs (ht = A ht-1 + B xt, ZOH discretization, input-dependent selection) inside the standard HuBERT pre-training loop, evaluated on SUPERB, phone purity, CCA, and long-context/streaming ASR. RS framework derives J-cost, φ-ladder, 8-tick periodicity, D=3, and constants from a single distinction via machine-checked theorems (e.g., reality_from_one_distinction, alexander_duality_circle_linking, costAlphaLog_fourth_deriv_at_zero). No shared structure, no parameter-free constant derivation, no 8-tick or J-cost reasoning appears; domain is applied speech SSL.","tokens_in":49072,"confidence":"high","tokens_out":287,"duration_ms":12129,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Mamba-based models can replace Transformers in HuBERT-style speech self-supervised learning and deliver lower compute for long audio plus stronger streaming results.","keywords":["Mamba","HuBERT","speech self-supervised learning","automatic speech recognition","streaming ASR","Selective State Space","quantized representations"],"falsifier":"Running the identical HuBERT pre-training schedule on a standard speech corpus and finding that the Mamba version scores lower than the Transformer version on a held-out long-context ASR test set.","tokens_in":2565,"feed_emoji":"🎙️","tokens_out":611,"duration_ms":21549,"temperature":0.7,"pith_summary":"The paper sets out to show that the Mamba architecture can be inserted into the standard HuBERT pre-training and fine-tuning pipeline for speech. A reader would care because the linear-time Selective State Space mechanism removes the quadratic cost that limits how much speech context Transformers can handle. The models match or beat Transformer baselines on SUPERB probes, produce cleaner quantized speech units, and separate speaker information more sharply. They also cut compute sharply when fine-tuned on long recordings and improve streaming ASR accuracy.","feed_headline":"Mamba models match Transformers on speech SSL with lower cost for long audio","feed_subtitle":"Linear-time state space lets them handle extended recordings and streaming tasks more efficiently while producing clearer speech units.","key_machinery":"The Selective State Space model inside Mamba, substituted directly into the HuBERT encoder stack to replace self-attention layers.","core_discovery":"Mamba-based HuBERT models achieve competitive performance on SUPERB benchmarks, especially in causal settings, while yielding higher-quality quantized representations and more distinct speaker features than Transformer counterparts. The linear-time property of the Selective State Space model lets the same pre-training recipe support fine-tuning on long-context ASR at much lower compute cost and produces superior results when the same models are adapted for streaming ASR.","pith_inferences":["The same linear scaling could let researchers pre-train on entire long-form audio files instead of fixed short segments.","Real-time speech systems might adopt these models for lower latency without sacrificing accuracy.","The clearer speaker separation could simplify downstream tasks such as diarization or voice conversion."],"forward_implications":["Fine-tuning for long-context ASR requires significantly lower compute than Transformer models.","Streaming ASR fine-tuning yields higher accuracy than the Transformer baseline.","Probing results remain competitive on SUPERB tasks and improve in causal configurations.","Quantized speech units extracted from the model are of higher quality.","Speaker-related information is captured more distinctly in the learned representations."],"fun_headline_variants":["Mamba HuBERT rivals Transformers in causal speech SSL settings","Mamba models support low cost streaming and long audio ASR","Mamba SSL models capture speaker features more distinctly","Mamba HuBERT enables fine tuning on long context with less compute"],"cache_read_input_tokens":64,"weakest_assumption_plain":"Mamba can be dropped straight into the existing HuBERT pre-training and fine-tuning recipe without any extra architectural changes or hyper-parameter retuning and still produce equal or better speech representations.","fun_headline_variants_meta":{"raw":{"variants":["Mamba HuBERT rivals Transformers in causal speech SSL settings","Mamba models support low cost streaming and long audio ASR","Mamba SSL models capture speaker features more distinctly","Mamba HuBERT enables fine tuning on long context with less compute"]},"model":"grok-4.3","cost_usd":0.011544,"raw_usage":{"total_tokens":4947,"prompt_tokens":605,"num_sources_used":0,"completion_tokens":65,"cost_in_usd_ticks":115440500,"prompt_tokens_details":{"text_tokens":605,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":4277,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":605,"tokens_out":65,"duration_ms":42949,"temperature":1.0,"reasoning_tokens":4277,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-19T09:10:46.239890+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running the identical HuBERT pre-training schedule on a standard speech corpus and finding that the Mamba version scores lower than the Transformer version on a held-out long-context ASR test set.","supporting_citations":[],"review_version":1}