REVIEW 4 major objections 5 minor 73 references
Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Dense feature prepending offers no clear quality edge over cross-attention for speech-to-text models.
desk verdict A clean, from-scratch comparison showing DFP does not beat cross-attention for speech-to-text at this scale, but the paper overstates 'parity' where the statistics only support 'no significant difference found.' read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The comparison rests on three architectural families with identical encoder/decoder layer counts and dimensions: cross-attention (a standard encoder-decoder with cross-attention in the decoder), decoder-prepend (a speech encoder's output is concatenated with text embeddings in a decoder-only stack), and decoder-only (no speech encoder; only a length adapter before prepending). Around these, the paper varies two mechanisms: CTC compression, which merges consecutive frames with identical predictions to shorten the audio sequence, and sequence-level knowledge distillation, which replaces target translations with more monotonic synthetic translations. A final manipulated variable is the causal mask over the concatenated speech+text sequence, which either restricts speech frames to look only backward or lets them attend to all speech frames.
What would settle it
Train decoder-prepend and cross-attention models at a much larger scale (for instance, 1B+ parameters) on the same data, with per-architecture hyperparameter search, and compare WER/BLEU and throughput. If dense feature prepending then clearly beats cross-attention in quality or efficiency, the paper's conclusion would be limited to small from-scratch models.
Extended reading notes
Core claim
The central claim is that, under controlled from-scratch training at the 65M–153M parameter scale, dense feature prepending (DFP) does not give a clear quality advantage over cross-attention for speech-to-text. In the Transformer setting, cross-attention is slightly better on average; in the Conformer-with-CTC setting, decoder-prepend and cross-attention are statistically indistinguishable on both ASR and ST. On efficiency, cross-attention is consistently faster and uses less GPU memory than decoder-prepend, while decoder-only models are slower, more memory-hungry, and lower in quality. Two further results qualify the picture: CTC compression is more beneficial for DFP than for cross-attention, and causal masking of the speech portion helps decoder-prepend but substantially hurts decoder-only models.
Load-bearing premise
The load-bearing premise is that the same training recipe—optimizer, learning-rate schedule, batch size, and early stopping—is a fair basis for comparing the three architectures; if each architecture needs different hyperparameters to perform well, the measured gaps could be artifacts.
Editorial extensions
If this is right
- Cross-attention is a strong, simple baseline for any future speech-LLM integration; DFP's adoption is not justified by quality at this scale.
- CTC compression is a cheap and effective technique for DFP models, giving most of the speed and memory benefit without quality loss.
- Causal masking should be kept on decoder-prepend models and removed for decoder-only models; a speech encoder changes how the decoder should mask audio.
- Decoder-only speech models at this scale are dominated by both cross-attention and decoder-prepend, which supports keeping a speech encoder when pursuing the LLM integration path.
Reading between the lines
- If the pattern holds at larger scale, many production speech-language systems built on DFP could switch to cross-attention without quality loss and gain throughput and memory headroom.
- The causality result suggests that in DFP the speech encoder is doing the full-attention work that the decoder's causal self-attention cannot; without an encoder, full attention over speech is necessary.
- The fixed training recipe leaves open that architecture-specific tuning could narrow or widen the gaps, so the conclusion is best read as 'DFP has no clear advantage under a common default recipe' rather than as a ceiling for DFP.
- A direct extension would be to test instruction-formatted prompts between speech and text, since the paper notes this LLM-specific modeling choice is absent from its setting and could alter the comparison.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a controlled empirical comparison of three ways to integrate speech into sequence-to-sequence text models: a standard encoder-decoder with cross-attention, a decoder-only model in which downsampled audio is prepended to text tokens, and a decoder-prepend model in which a speech encoder's representations are prepended. All models are trained from scratch on MuST-C v1.0 and CoVoST2, with Transformer and Conformer encoders, and are evaluated on ASR and ST across monolingual, bilingual, and multilingual settings, including CTC auxiliary loss and compression, sequence-level knowledge distillation, generation speed, GPU memory, and a causal-masking ablation. The main empirical conclusion is that, at the tested scale, decoder-prepend does not show a clear advantage over cross-attention, both outperform decoder-only models on quality and efficiency, and CTC compression is more favorable to decoder-prepend than to cross-attention.
Significance. The study is a timely and useful counterpoint to the common assumption that dense feature prepending is the preferred way to connect speech encoders to LLMs. Its strengths include a deliberately controlled setup (same data, same training recipe, models trained from scratch), evaluation over 12 language directions and two tasks, multiple configurations (Transformer/Conformer, CTC compression, seqKD), significance tests for the main comparisons, and released code under a permissive license. If the central negative claim is accepted, the paper provides practical guidance that cross-attention is a safe, more efficient default at this scale. The main fragility is inferential: the evidence is stronger for 'we did not find an advantage' than for 'the architectures are on par,' and the paper's own wording alternates between the two; this needs to be resolved in revision.
major comments (4)
- [§5.1, Table 1 (lines 4–5)] The sentence 'None of these differences are statistically significant' is immediately followed by 'The above results show that decoder-prepend is on par with cross-attention.' This is not a valid equivalence inference: a null result from a significance test only means the test did not detect a difference, and with single runs it says nothing about the size of the true difference. The data could be consistent with a meaningfully worse or better decoder-prepend under a different sample or seed. To support the parity language in the Introduction ('overall similar results') and §5.1, the authors should add pre-specified equivalence margins or confidence intervals for the WER/BLEU differences; otherwise the conclusions should be rephrased as absence-of-evidence claims, consistent with the more cautious Abstract wording.
- [§4.3] All three architectures are trained with one shared recipe (Adam, peak learning rate 2e-3, 25k warmup, batch sizes 320k/256k frames, early stopping patience 10) and there is no per-architecture hyperparameter tuning. A shared recipe is a reasonable control, but for a comparative negative claim the absence of any sensitivity analysis leaves open the possibility that the architecture ranking would change under different hyperparameters. Please report a small robustness check (e.g., learning rate or warmup variation) for the central comparison, or explicitly discuss the risk and the evidence that the chosen budget is fair to all three architectures.
- [§4.3, Table 1] Most reported results are single training runs, and the significance tests (bootstrap resampling for ASR, approximate randomization for ST) quantify only test-set sampling variability, not training stochasticity. The p-values and †/‡ markers in Table 1 are therefore insufficient to support statements such as 'None of these differences are statistically significant' as evidence about architecture superiority. The authors should either provide multiple seeds for the key cross-attention vs. decoder-prepend comparisons (at least Table 1 lines 4–5) or explicitly state that the significance tests ignore training variance.
- [§5.2, Table 1 (lines 4.1 and 5.1)] The claim that 'decoder-prepend better leverages CTC compression' is based on the observation that compression degrades cross-attention (WER 19.6→21.8, BLEU 29.7→28.5) while decoder-prepend is unchanged (19.9→19.9, 29.7→29.7). This is a difference-in-differences claim, but no interaction test or confidence intervals are provided for those differences, and both comparisons rest on single runs. Please test the architecture-by-compression interaction directly or provide intervals, so readers can judge whether the asymmetry is beyond noise.
minor comments (5)
- [Table 1] In the plain-text rendering, the #Params entry for 'decoder-only 18L' is not visible; please ensure every row reports parameter counts and that the table is not missing a cell.
- [§4.3] 'Experiments are run on 4 Nvidia A100-40GB GPUs for about 2 days' is ambiguous about whether this is per model or for the whole suite; please state the per-model training time.
- [Figure 2 and §5.2] The abbreviations 'compr' and 'CF-compr' are used inhomogeneously; define both in one place and use them consistently.
- [§5.3] The claim that seqKD makes the target 'more monotonically aligned' to the source should cite the specific analysis in Zhou et al. (2019), not only the general seqKD method.
- [§5.5] The 8.18 BLEU drop for decoder-prepend TF on multilingual ST when causal masking is removed is striking and is attributed to longer inputs; a per-language breakdown or an analysis by input length would make the hypothesis testable.
Circularity Check
No significant circularity: the comparison is an empirical benchmark study whose conclusions rest on measured WER/BLEU values, not on fitted parameters or self-citation chains.
full rationale
This paper is an empirical comparison of three architectures (cross-attention, decoder-only, and decoder-prepend) on standard public benchmarks (MuST-C v1.0, CoVoST2). The central claim, that DFP shows no clear advantage over cross-attention, is supported by reported WER and BLEU numbers in Tables 1–3. No step in the derivation chain reduces to its own inputs: the architectures are defined by their attention mechanisms (Section 3), all models are trained from scratch with the same recipe (Section 4.3), and the results are measurements on held-out test sets. The self-citations to Gaido et al. (2021) for CTC compression and Papi et al. (2024a) for the Conformer implementation are tooling references, not load-bearing assumptions that predetermine the comparison outcome. The limitation that only p>0.05 is used to support parity is a statistical inference concern, not a circularity: non-significance does not equal equivalence, but this does not mean the paper's results were constructed from its conclusions. There is no fitting-to-target, no renaming of a known result as a new derivation, and no imported uniqueness theorem. Accordingly, the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (6)
- CTC auxiliary loss weight =
0.5
- Peak learning rate =
2e-3
- Batch size in audio frames =
320k (MuST-C), 256k (CoVoST2)
- Early stopping patience =
10
- Layer distribution =
12 encoder / 6 decoder, or 18 decoder-only, or 32 decoder-only
- SpecAugment mask sizes =
frequency 27, time 100
assumptions (4)
- domain assumption The same training recipe (learning rate, batch size, warmup, early stopping) is equally suitable for all three architectures.
- domain assumption Models of 65 to 153 million parameters trained from scratch are representative enough to compare architectural design choices that also matter at LLM scale.
- domain assumption MuST-C v1.0 and CoVoST2 are adequate and representative benchmarks for comparing S2T architectures.
- domain assumption WER and sacreBLEU are sufficient quality metrics for the architectural comparison.
Cite this review
Pith. "Pith review of Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison." pith.science (2026). https://pith.science/paper/W3UVL4ER
@misc{pith2026250102370,
author = {Pith},
title = {Pith review of: Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison},
year = {2026},
howpublished = {\url{https://pith.science/paper/W3UVL4ER}},
note = {Machine review of arXiv:2501.02370}
}
read the original abstract
Following the remarkable success of Large Language Models (LLMs) in NLP tasks, there is increasing interest in extending their capabilities to speech -- the most common form of communication. The most widespread approach to integrating speech into LLMs is dense feature prepending (DFP), which prepends the projected speech representations to the textual representations, allowing end-to-end training with a speech encoder. This raises questions about the need for a sophisticated speech encoder for DFP and how its performance compares with a standard encoder-decoder (i.e., cross-attention) architecture. We compare DFP and cross-attention under a variety of configurations, such as CTC compression, sequence-level knowledge distillation, on monolingual, bilingual, and multilingual models. To perform a controlled architectural comparison, we train all models from scratch rather than using large pretrained models and use comparable data and parameter settings, testing speech-to-text recognition (ASR) and translation (ST) on MuST-C v1.0 and CoVoST2 datasets. Despite the wide adoption of DFP, our results do not indicate a clear advantage of DFP over cross-attention.
Figures
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Ibrahim Said Ahmad, Antonios Anastasopoulos, Ond r ej Bojar, Claudia Borg, Marine Carpuat, Roldano Cattoni, Mauro Cettolo, William Chen, Qianqian Dong, Marcello Federico, Barry Haddow, D \'a vid Javorsk \'y , Mateusz Krubi \'n ski, Tsz Kin Lam, Xutai Ma, Prashant Mathur, Evgeny Matusov, Chandresh Maurya, John McCrae, Kenton Murray, Satoshi Nakamura, Matte...
2024
-
[4]
Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, and Furu Wei. 2021. https://arxiv.org/abs/2110.07205 Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing
arXiv 2021
-
[5]
Sameer Bansal, Herman Kamper, Adam Lopez, and Sharon Goldwater. 2017. https://aclanthology.org/E17-2076/ Towards speech-to-text translation without speech recognition . In Proceedings of the 15th Conference of the E uropean Chapter of the Association for Computational Linguistics: Volume 2, Short Papers , pages 474--479, Valencia, Spain. Association for C...
work page 2017
-
[6]
Lo \" c Barrault, Yu-An Chung, Mariano Cora Meglioli, David Dale, Ning Dong, Paul-Ambroise Duquenne, Hady Elsahar, Hongyu Gong, Kevin Heffernan, John Hoffman, et al. 2023. Seamlessm4t-massively multilingual & multimodal machine translation. arXiv preprint arXiv:2308.11596
arXiv 2023
-
[7]
Alexandre B \'e rard, Laurent Besacier, Ali Can Kocabiyikoglu, and Olivier Pietquin. 2018. End-to-End Automatic Speech Translation of Audiobooks . In Proceedings of ICASSP 2018 - IEEE International Conference on Acoustics, Speech and Signal Processing , Calgary, Alberta, Canada
work page 2018
-
[8]
Alexandre B \'e rard, Olivier Pietquin, Christophe Servan, and Laurent Besacier. 2016. Listen and translate: A proof of concept for end-to-end speech-to-text translation. arXiv preprint arXiv:1612.01744
arXiv 2016
Show all 73 references
-
[9]
Tom B Brown. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165
2020 arXiv
-
[10]
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals. 2016. https://doi.org/10.1109/ICASSP.2016.7472621 Listen, attend and spell: A neural network for large vocabulary conversational speech recognition . In 2016 IEEE International Conference on Acoustics, Speech and Signal...
2016
-
[11]
William Chan, Navdeep Jaitly, Quoc V Le, and Oriol Vinyals. 2015. Listen, attend and spell. arXiv preprint arXiv:1508.01211
2015 arXiv
-
[12]
Feilong Chen, Minglun Han, Haozhi Zhao, Qingyang Zhang, Jing Shi, Shuang Xu, and Bo Xu. 2023. X-llm: Bootstrapping advanced large language models by treating multi-modalities as foreign languages. arXiv preprint arXiv:2305.04160
2023 arXiv
-
[13]
Xi Chen, Songyang Zhang, Qibing Bai, Kai Chen, and Satoshi Nakamura. 2024 a . https://doi.org/10.18653/v1/2024.findings-acl.416 LL a ST : Improved end-to-end speech translation system leveraged by large language models . In Findings of the Association for Computational Linguis...
2024 doi
-
[14]
Zhehuai Chen, He Huang, Andrei Andrusenko, Oleksii Hrinchuk, Krishna C Puvvada, Jason Li, Subhankar Ghosh, Jagadeesh Balam, and Boris Ginsburg. 2024 b . Salm: Speech-augmented language model with in-context learning for speech recognition and translation. In ICASSP 2024-2024 I...
2024
-
[15]
Puvvada, Nithin Rao Koluguri, Piotr Żelasko, Jagadeesh Balam, and Boris Ginsburg
Zhehuai Chen, He Huang, Oleksii Hrinchuk, Krishna C. Puvvada, Nithin Rao Koluguri, Piotr Żelasko, Jagadeesh Balam, and Boris Ginsburg. 2024 c . https://arxiv.org/abs/2406.19954 Bestow: Efficient and streamable speech language model with the best of two worlds in gpt and t5 . P...
2024 arXiv
-
[16]
Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shiliang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou. 2023. Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models. arXiv preprint arXiv:2311.07919
2023 arXiv
-
[17]
Marta R Costa-juss \`a , James Cross, Onur C elebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. 2022. No language left behind: Scaling human-centered machine translation. arXiv preprint arXiv:2207.04672
2022 arXiv
-
[18]
Di Gangi, Roldano Cattoni, Luisa Bentivogli, Matteo Negri, and Marco Turchi
Mattia A. Di Gangi, Roldano Cattoni, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2019 a . https://doi.org/10.18653/v1/N19-1202 M u ST - C : a M ultilingual S peech T ranslation C orpus . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...
2019 doi
-
[19]
Mattia Antonino Di Gangi, Matteo Negri, Roldano Cattoni, Roberto Dessi, and Marco Turchi. 2019 b . https://aclanthology.org/W19-6603/ Enhancing transformer for end-to-end speech-to-text translation . In Proceedings of Machine Translation Summit XVII: Research Track, pages 21--...
2019
-
[20]
Qingkai Fang, Rong Ye, Lei Li, Yang Feng, and Mingxuan Wang. 2022. https://doi.org/10.18653/v1/2022.acl-long.486 STEMM : Self-learning with speech-text manifold mixup for speech translation . In Proceedings of the 60th Annual Meeting of the Association for Computational Lingui...
2022 doi
-
[21]
Yassir Fathullah, Chunyang Wu, Egor Lakomkin, Junteng Jia, Yuan Shangguan, Ke Li, Jinxi Guo, Wenhan Xiong, Jay Mahadeokar, Ozlem Kalinli, et al. 2024. Prompting large language models with speech recognition abilities. In ICASSP 2024-2024 IEEE International Conference on Acoust...
2024
-
[22]
Marco Gaido, Mauro Cettolo, Matteo Negri, and Marco Turchi. 2021. https://doi.org/10.18653/v1/2021.eacl-main.57 CTC -based compression for direct speech translation . In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics...
2021 doi
-
[23]
Marco Gaido, Sara Papi, Dennis Fucci, Giuseppe Fiameni, Matteo Negri, and Marco Turchi. 2022. https://doi.org/10.18653/v1/2022.iwslt-1.13 Efficient yet competitive speech translation: FBK @ IWSLT 2022 . In Proceedings of the 19th International Conference on Spoken Language Tra...
2022 doi
-
[24]
Marco Gaido, Sara Papi, Matteo Negri, and Luisa Bentivogli. 2024. https://doi.org/10.18653/v1/2024.acl-long.789 Speech translation with speech foundation models and large language models: What is there and what is missing? In Proceedings of the 62nd Annual Meeting of the Assoc...
2024 doi
-
[25]
Alex Graves, Santiago Fern \'a ndez, Faustino Gomez, and J \"u rgen Schmidhuber. 2006. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd international conference on Machine learning, pages 369--376
2006
-
[26]
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang. 2020. https://doi.org/10.21437/Interspeech.2020-3015 Conformer: Convolution-augmented transformer for speech recognition . In Inters...
2020 doi
-
[27]
Ankit Gupta, George Saon, and Brian Kingsbury. 2024. https://doi.org/10.21437/Interspeech.2024-565 Exploring the limits of decoder-only models trained on public speech recognition corpora . In Interspeech 2024, pages 252--256
2024 doi
-
[28]
Yukiya Hono, Koh Mitsuda, Tianyu Zhao, Kentaro Mitsui, Toshiaki Wakatsuki, and Kei Sawada. 2024. https://doi.org/10.18653/v1/2024.findings-acl.787 Integrating pre-trained speech and language models for end-to-end speech recognition . In Findings of the Association for Computat...
2024 doi
-
[29]
Shujie Hu, Long Zhou, Shujie Liu, Sanyuan Chen, Hongkun Hao, Jing Pan, Xunying Liu, Jinyu Li, Sunit Sivasankaran, Linquan Liu, et al. 2024. Wavllm: Towards robust and adaptive speech large language model. arXiv preprint arXiv:2404.00656
2024 arXiv
-
[30]
Chao-Wei Huang, Hui Lu, Hongyu Gong, Hirofumi Inaguma, Ilia Kulikov, Ruslan Mavlyutov, and Sravya Popuri. 2024 a . Investigating decoder-only large language models for speech-to-text translation. arXiv preprint arXiv:2407.03169
2024 arXiv
-
[31]
Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, et al. 2024 b . Audiogpt: Understanding and generating speech, music, sound, and talking head. In Proceedings of the AAAI Conference on Artifici...
2024
-
[32]
Hirofumi Inaguma, Tatsuya Kawahara, and Shinji Watanabe. 2021. https://doi.org/10.18653/v1/2021.naacl-main.150 Source and target bidirectional knowledge distillation for end-to-end speech translation . In Proceedings of the 2021 Conference of the North American Chapter of the ...
2021 doi
-
[33]
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825
2023 arXiv
-
[34]
Yoon Kim and Alexander M. Rush. 2016. https://doi.org/10.18653/v1/D16-1139 Sequence-level knowledge distillation . In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1317--1327, Austin, Texas. Association for Computational Linguistics
2016 doi
-
[35]
Philipp Koehn. 2004. https://aclanthology.org/W04-3250/ Statistical significance tests for machine translation evaluation . In Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing, pages 388--395, Barcelona, Spain. Association for Computationa...
2004
-
[36]
Taku Kudo and John Richardson. 2018. https://doi.org/10.18653/v1/D18-2012 S entence P iece: A simple and language independent subword tokenizer and detokenizer for neural text processing . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processin...
2018 doi
-
[37]
Egor Lakomkin, Chunyang Wu, Yassir Fathullah, Ozlem Kalinli, Michael L Seltzer, and Christian Fuegen. 2024. End-to-end speech recognition contextualization with large language models. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing ...
2024
-
[38]
Tsz Kin Lam, Shigehiko Schamoni, and Stefan Riezler. 2022. https://doi.org/10.18653/v1/2022.acl-short.27 Sample, translate, recombine: Leveraging audio alignments for data augmentation in end-to-end speech translation . In Proceedings of the 60th Annual Meeting of the Associat...
2022 doi
-
[39]
Tsz Kin Lam, Shigehiko Schamoni, and Stefan Riezler. 2023. Make more of your data: Minimal effort data augmentation for automatic speech recognition and translation. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1--5. IEEE
2023
-
[40]
Siddique Latif, Moazzam Shoukat, Fahad Shamshad, Muhammad Usama, Heriberto Cuay \'a huitl, and Bj \"o rn W Schuller. 2023. Sparks of Large Audio Models: A Survey and Outlook . arXiv preprint arXiv:2308.12792
2023 arXiv
-
[41]
Beomseok Lee, Ioan Calapodescu, Marco Gaido, Matteo Negri, and Laurent Besacier. 2024. Speech-massive: A multilingual speech dataset for slu and beyond. arXiv preprint arXiv:2408.03900
2024 arXiv
-
[42]
Jinyu Li et al. 2022. Recent advances in end-to-end automatic speech recognition. APSIPA Transactions on Signal and Information Processing, 11(1)
2022
-
[43]
Yuchen Liu, Junnan Zhu, Jiajun Zhang, and Chengqing Zong. 2020. Bridging the modality gap for speech-to-text translation. arXiv preprint arXiv:2010.14920
2020 arXiv
-
[44]
Cosmin Munteanu, Matt Jones, Sharon Oviatt, Stephen Brewster, Gerald Penn, Steve Whittaker, Nitendra Rajput, and Amit Nanavati. 2013. https://doi.org/10.1145/2468356.2468803 We need to talk: HCI and the delicate topic of spoken language interaction . In CHI '13 Extended Abstra...
2013
-
[45]
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. https://doi.org/10.18653/v1/N19-4009 fairseq: A fast, extensible toolkit for sequence modeling . In Proceedings of the 2019 Conference of the North A merican Chap...
2019 doi
-
[46]
Jing Pan, Jian Wu, Yashesh Gaur, Sunit Sivasankaran, Zhuo Chen, Shujie Liu, and Jinyu Li. 2023. Cosmic: Data efficient instruction-tuning for speech in-context learning. arXiv preprint arXiv:2311.02248
2023 arXiv
-
[47]
Sara Papi, Marco Gaido, Andrea Pilzer, and Matteo Negri. 2024 a . https://doi.org/10.18653/v1/2024.acl-long.200 When good and reproducible results are a giant with feet of clay: The importance of software quality in NLP . In Proceedings of the 62nd Annual Meeting of the Associ...
2024 doi
-
[48]
Sara Papi, Peter Polak, Ond r ej Bojar, and Dominik Mach \'a c ek. 2024 b . How" real" is your real-time simultaneous speech-to-text translation system? arXiv preprint arXiv:2412.18495
2024 arXiv
-
[49]
Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le. 2019. https://doi.org/10.21437/Interspeech.2019-2680 Specaugment: A simple data augmentation method for automatic speech recognition . In Interspeech 2019, pages 2613--2617
2019 doi
-
[50]
Matt Post. 2018. https://doi.org/10.18653/v1/W18-6319 A call for clarity in reporting BLEU scores . In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186--191, Brussels, Belgium. Association for Computational Linguistics
2018 doi
-
[51]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...
2021
-
[52]
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In International conference on machine learning, pages 28492--28518. PMLR
2023
-
[53]
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners
2019
-
[54]
Stefan Riezler and John T. Maxwell. 2005. https://aclanthology.org/W05-0908/ On some pitfalls in automatic evaluation and significance testing for MT . In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarizatio...
2005
-
[55]
Matthias Sperber and Matthias Paulik. 2020. https://doi.org/10.18653/v1/2020.acl-main.661 Speech translation and the end-to-end promise: Taking stock of where we are . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7409--7421,...
2020 doi
-
[56]
Chameleon Team. 2024. Chameleon: Mixed-modal early-fusion foundation models. arXiv preprint arXiv:2405.09818
2024 arXiv
-
[57]
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivi \`e re, Mihir Sanjay Kale, Juliette Love, et al. 2024. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295
2024 arXiv
-
[58]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288
2023 arXiv
-
[59]
G \'a llego, Jos \'e A
Ioannis Tsiamas, Gerard I. G \'a llego, Jos \'e A. R. Fonollosa, and Marta R. Costa-juss \`a . 2024. https://doi.org/10.18653/v1/2024.findings-acl.847 Pushing the limits of zero-shot end-to-end speech translation . In Findings of the Association for Computational Linguistics: ...
2024 doi
-
[60]
Gállego, José A
Ioannis Tsiamas, Gerard I. Gállego, José A. R. Fonollosa, and Marta R. Costa-jussà. 2022. https://doi.org/10.21437/Interspeech.2022-59 Shas: Approaching optimal segmentation for end-to-end speech translation . In Interspeech 2022, pages 106--110
2022 doi
-
[61]
Emiru Tsunoo, Hayato Futami, Yosuke Kashiwagi, Siddhant Arora, and Shinji Watanabe. 2023. Decoder-only architecture for speech recognition with ctc prompts and text data augmentation. arXiv preprint arXiv:2309.08876
2023 arXiv
-
[62]
Emiru Tsunoo, Hayato Futami, Yosuke Kashiwagi, Siddhant Arora, and Shinji Watanabe. 2024. Decoder-only architecture for streaming end-to-end speech recognition. arXiv preprint arXiv:2406.16107
2024 arXiv
-
[63]
A Vaswani. 2017. Attention is all you need. Advances in Neural Information Processing Systems
2017
-
[64]
Francesco Verdini, Pierfrancesco Melucci, Stefano Perna, Francesco Cariaggi, Marco Gaido, Sara Papi, Szymon Mazurek, Marek Kasztelnik, Luisa Bentivogli, Sébastien Bratières, Paolo Merialdo, and Simone Scardapane. 2024. https://arxiv.org/abs/2409.17044 How to connect speech fou...
2024 arXiv
-
[65]
Changhan Wang, Yun Tang, Xutai Ma, Anne Wu, Dmytro Okhonko, and Juan Pino. 2020. https://doi.org/10.18653/v1/2020.aacl-demo.6 Fairseq S 2 T : Fast speech-to-text modeling with fairseq . In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Com...
2020 doi
-
[66]
Changhan Wang, Anne Wu, Jiatao Gu, and Juan Pino. 2021. Covost 2 and massively multilingual speech translation. In Interspeech, pages 2247--2251
2021
-
[67]
Mingqiu Wang, Wei Han, Izhak Shafran, Zelin Wu, Chung-Cheng Chiu, Yuan Cao, Yongqiang Wang, Nanxin Chen, Yu Zhang, Hagen Soltau, et al. 2023. Slm: Bridge the thin gap between speech and text foundation models. arXiv preprint arXiv:2310.00230
2023 arXiv
-
[68]
Ron J Weiss, Jan Chorowski, Navdeep Jaitly, Yonghui Wu, and Zhifeng Chen. 2017. Sequence-to-sequence models can directly translate foreign speech. arXiv preprint arXiv:1703.08581
2017 arXiv
-
[69]
Jian Wu, Yashesh Gaur, Zhuo Chen, Long Zhou, Yimeng Zhu, Tianrui Wang, Jinyu Li, Shujie Liu, Bo Ren, Linquan Liu, and Yu Wu. 2023. https://doi.org/10.1109/ASRU57964.2023.10389705 On decoder-only architecture for speech-to-text and large language model integration . In 2023 IEE...
2023
-
[70]
Chenyu You, Nuo Chen, Fenglin Liu, Shen Ge, Xian Wu, and Yuexian Zou. 2022. https://doi.org/10.18653/v1/2022.findings-naacl.91 End-to-end spoken conversational question answering: Task, dataset and model . In Findings of the Association for Computational Linguistics: NAACL 202...
2022 doi
-
[71]
Piotr \.Z elasko, Zhehuai Chen, Mengru Wang, Daniel Galvez, Oleksii Hrinchuk, Shuoyang Ding, Ke Hu, Jagadeesh Balam, Vitaly Lavrukhin, and Boris Ginsburg. 2024. Emmett: Efficient multimodal machine translation training. arXiv preprint arXiv:2409.13523
2024 arXiv
-
[72]
Chunting Zhou, Graham Neubig, and Jiatao Gu. 2019. Understanding knowledge distillation in non-autoregressive machine translation. arXiv preprint arXiv:1911.02727
2019 arXiv
-
[73]
Maike Z \"u fle and Jan Niehues. 2024. Contrastive learning for task-independent speechllm-pretraining. arXiv preprint arXiv:2412.15712
2024 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.