REVIEW 4 major objections 4 minor 1 cited by
AMuSeD: An Attentive Deep Neural Network for Multimodal Sarcasm Detection Incorporating Bi-modal Data Augmentation
T0 review · 4 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper claims that a text-audio sarcasm detector trained on back-translated text and synthesized speech reaches an F1-score of 81.0 percent on the MUStARD dataset, beating models that also use video.
desk verdict The augmentation idea is worth a look, but the 81% F1 can't be trusted until the authors rule out test-fold leakage in their data generation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the bimodal augmentation pipeline: text is back-translated as $t^L_b = \mathrm{BackTranslation}(t_o)$ through Greek, German, French, and Italian, and aligned audio $a^L$ is generated from each augmented text by Amazon Polly or FastSpeech 2, with FastSpeech 2 fine-tuned on MUStARD so that sarcastic intonation survives synthesis. On top of that, text is encoded by BERT and audio by VGGish, and the two feature matrices are fused by applying self-attention separately to each modality and then adding skip connections before concatenation, so the model keeps the original context while focusing on sarcasm-relevant features.
What would settle it
Rerun the 20-fold experiment with augmented samples and the fine-tuned TTS model generated exclusively from training-fold utterances; if the F1-score falls substantially below the reported 81.0 percent, the gap would show that the result relies on test-set information entering through the augmentation pipeline.
Extended reading notes
Core claim
On the MUStARD dataset, the paper's central claim is that bimodal text-audio augmentation with self-attention fusion gives an F1-score of 81.0 percent, surpassing models that additionally use the video modality. The authors also claim that larger augmented datasets improve performance, that the Fine-tuned FastSpeech 2 synthesizer produces speech judged closer to sarcasm than Amazon Polly or pre-trained FastSpeech 2, and that self-attention with skip connections is the best fusion choice among the attention mechanisms they test. In their comparison, the best configuration uses 20-fold augmented data built from 10,460 audio samples and 11,150 text samples.
Load-bearing premise
The load-bearing premise is that the offline augmentation pipeline and the fine-tuned FastSpeech 2 model are built using only training-fold utterances; the paper describes augmenting all 690 utterances and never states that the test fold is excluded from augmentation or TTS fine-tuning.
Editorial extensions
If this is right
- Holding the synthesizer fixed, F1 rises with augmentation volume: the 20-fold dataset reaches 80.98 percent, beating 16-fold and 4-fold versions.
- On the 4-fold dataset, Fine-tuned FastSpeech 2 outperforms both Amazon Polly and pre-trained FastSpeech 2, and listeners rate its audio higher on quality and sarcasm resemblance.
- Self-attention with skip connections improves F1 over self-attention alone, while cross-attention loses performance when skip connections are added.
- BERT text features beat GloVe on both original and augmented data, and text-audio fusion beats either modality alone on the unaugmented baseline.
Reading between the lines
- A direct extension would be to run the same pipeline on other small multimodal sarcasm or emotion datasets to see whether the gain is specific to MUStARD or generalizes.
- Because the method replaces video with synthetic prosody, a natural ablation would classify using synthesized audio alone versus real audio alone to isolate how much of the gain comes from TTS quality.
- A speaker-independent evaluation, where no speaker appears in both training and test folds, would separate the augmentation's effect from speaker-specific leakage in the MUStARD split.
- Future work implied by the paper is to add a video branch on top of the same augmentation and measure whether synthetic audio still contributes once real facial cues are available.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes AMuSeD, a text-audio sarcasm detector trained on the MUStARD dataset with offline bimodal augmentation: back-translation of all 690 texts through four pivot languages, and audio synthesis using Amazon Polly, pre-trained FastSpeech 2, and FastSpeech 2 fine-tuned on MUStARD. Text and audio features are extracted with BERT and VGGish, fused by a self-attention module with a skip connection, and classified by a fully connected layer. The paper reports an F1-score of 81.0% under the official 5-fold MUStARD protocol, surpassing trimodal baselines, and a listening test indicating that fine-tuned FastSpeech 2 audio is perceived as higher quality and more sarcastic. The experimental narrative is internally consistent, since the final F1 appears in both Tables V and VII (80.98/81.0), but the evaluation protocol does not yet establish the central claim because the augmentation and TTS fine-tuning procedures are described at the dataset level rather than at the fold level.
Significance. AMuSeD addresses a real problem: MUStARD contains only about 690 utterances, which is very small for deep multimodal models. The idea of coupling back-translated text with sarcasm-tuned text-to-speech is a useful and potentially transferable contribution if it can be shown that the augmentation pipeline does not leak test information. The architecture is transparent and reproducible in principle, using standard encoders (BERT, VGGish), a simple fusion module, and standard training details. The listening test is a welcome addition to the paper. However, the scientific value is conditional on the leakage issue: if test-fold utterances entered back-translation or FastSpeech 2 fine-tuning, then the reported 81.0% F1 and the claimed superiority over trimodal models are not evidence for the proposed method. The paper would also be materially strengthened by fold-wise variance estimates, since all headline comparisons are point estimates.
major comments (4)
- [Section III-A Steps 1-3, Table III, Section IV-A-3] The augmentation pipeline is not described as fold-isolated. Step 1 states that the initial training dataset consists of 690 samples, Step 2 states that FastSpeech 2 was fine-tuned on the MUStARD dataset, and Table III builds augmented datasets from all 690 utterances. The evaluation in Section IV-A-3 uses the official 5-fold split, but no sentence anywhere restricts back-translation, audio synthesis, or FastSpeech 2 fine-tuning to the training portion of each fold. If test-fold texts are back-translated, their paraphrased variants can appear in the training set; if test-fold audio is used to fine-tune FastSpeech 2, the synthesized training audio can encode test-set prosody. Both channels would directly inflate the Table V F1 of 81.0 and invalidate the comparison with the baselines. The authors must re-run the entire augmentation and TTS fine-tuning pipeline inside each training fold, or provide explicit evidence that no test-fold information was used in any augmentation stage.
- [Section IV-B-3, Tables VI and VII] The choice of self-attention as the final fusion mechanism appears to be made after evaluating several attention variants on the test folds, with no held-out validation split described. Table VI reports that cross-attention achieves F1 80.88, which is higher than self-attention without skip connections (79.64); the final preference for self-attention rests on the additional effect of skip connections in Table VII. Because the selection among attention mechanisms is not described as a pre-specified or validation-based procedure, the reported 81.0% may overstate the expected performance of the chosen configuration. The authors should either define a validation-based model selection protocol or report all configurations with fold-wise variability.
- [Section IV-B-1, Figures 4-6, Tables V-VIII] All results are reported as point estimates without error bars, standard deviations, or significance tests. The claim that 81.0% is a significant improvement over the 76.7% of GEMA cannot be assessed from a single 5-fold estimate, and the internal comparisons in Figures 4 and 5 lack any measure of uncertainty. The authors should report per-fold scores or repeated-run statistics, and use a paired significance test when comparing models on the same folds.
- [Section III-C, Eq. (12)] Equation (12) is dimensionally inconsistent: it defines the skip-connected feature as a sum over i of the dot product M_i · m_tilde, which produces a scalar, while the text states that the attended textual vector t_tilde_s belongs to R^512 and is later concatenated into a 1024-dimensional vector in Eq. (13). This cannot be the operation used in the experiments. The authors should provide a corrected formulation, for example a residual connection such as m_tilde + M, and ensure that all stated dimensions are consistent.
minor comments (4)
- [Section III-A Step 3] The heading contains a typo: 'Text-audio Biomodal Data Augmentation' should read 'Bimodal'.
- [Section III-A Step 1] The sentence 'the first phase involves translating the original text to into a secondary language' contains a duplicated preposition; please correct it.
- [Table IV] The batch size is listed as the set [16, 32, 64, 128, 256] rather than a single value; please specify the selected batch size or describe how it was tuned.
- [Figure 6] The listening test is reported only as stacked percentage distributions; reporting mean opinion scores with confidence intervals or a significance test would make the claims about Fine-tuned FastSpeech 2 easier to evaluate.
Circularity Check
Reported 81.0% F1 is partially circular: FastSpeech 2 is fine-tuned on the full 690-utterance MUStARD corpus, then used to synthesize augmented audio for the same corpus on which 5-fold cross-validation is reported, so test-fold information enters the augmentation pipeline.
-
fitted input called prediction
[Section III-A Step 2 (Speech Augmentation) and Eq. (2c); Section IV-A-4 (evaluation protocol)]
"We first pre-trained FS2 on a diverse dataset, such as LibriTTS, to generate high-quality speech. We then fine-tuned it on the MUStARD dataset to enhance its capability in conveying sarcastic features. ... Our initial training dataset consists of 690 samples. ... In alignment with the evaluation protocol established by Castro et al. [8], our experiment uses a 5-fold cross-validation method. The dataset is split according to a predefined file accessible on GitHub."
The Fine-tuned FS2 audio synthesizer is a fitted model, and it is fitted once on the entire MUStARD corpus. The 5-fold protocol then evaluates AMuSeD on folds drawn from that same corpus. Eq. (2c) generates the augmented audio a^L_FS = FastSpeech2(t^L_b) from back-translated text with this same fine-tuned FS2, and Table III lists augmented audio built from the full 690 utterances. Therefore, for any utterance in a test fold, its synthetic counterpart in the augmented training data is generated by a TTS whose weights were optimized on that utterance's original audio; back-translated text versions of test utterances can also occur in training.
full rationale
The paper's central claim is an empirical benchmark result, not an analytic derivation, so most of the high-end circularity categories do not apply. The load-bearing circular element is the offline augmentation pipeline: Section III-A Step 1 says the initial dataset is the full 690 samples and Step 2 says FS2 is fine-tuned on the MUStARD dataset, with no statement restricting back-translation, synthesis, or fine-tuning to the training folds used in Section IV-A-4. Because Eq. (2c) uses this Fine-tuned FS2 to create the augmented audio that AMuSeD is trained on, test-fold acoustic information leaks into the training generator, and the reported F1 of 81.0% is partially a reconstruction of data the model has already seen. This is a fitted-input/called-prediction circularity rather than a self-citation chain. Self-citations such as [27] and [24] appear in the related-work and feature-extraction motivation, but they are not load-bearing for the benchmark claim, so they do not raise the score beyond the leakage issue. The attention-mechanism comparison (Tables VI-VII) is model selection on the test set, which is overfitting rather than circularity, and the same leakage concern applies to the listening test for fine-tuned FS2 if the 10 natural utterances were part of the fine-tuning data.
Assumptions & free parameters
free parameters (6)
- Augmentation scale (20-fold) =
10,460 audio samples
- Batch size =
16 to 256 (searched)
- Dense layer size and dropout =
512 and 0.5
- Token sequence length S =
20
- FastSpeech 2 fine-tuning iterations =
200,000
- Secondary languages for back translation =
Greek, German, French, Italian
assumptions (4)
- domain assumption Back-translated text preserves the original sarcasm label.
- domain assumption Synthesized speech retains sarcastic prosody well enough to be a valid training signal.
- domain assumption Augmented samples and TTS fine-tuning do not leak test-fold information.
- domain assumption Frozen BERT and VGGish features are sufficient for the task.
Cite this review
Pith. "Pith review of AMuSeD: An Attentive Deep Neural Network for Multimodal Sarcasm Detection Incorporating Bi-modal Data Augmentation." pith.science (2026). https://pith.science/paper/C4W6UDF3
@misc{pith2026241210103,
author = {Pith},
title = {Pith review of: AMuSeD: An Attentive Deep Neural Network for Multimodal Sarcasm Detection Incorporating Bi-modal Data Augmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/C4W6UDF3}},
note = {Machine review of arXiv:2412.10103}
}
read the original abstract
Detecting sarcasm effectively requires a nuanced understanding of context, including vocal tones and facial expressions. The progression towards multimodal computational methods in sarcasm detection, however, faces challenges due to the scarcity of data. To address this, we present AMuSeD (Attentive deep neural network for MUltimodal Sarcasm dEtection incorporating bi-modal Data augmentation). This approach utilizes the Multimodal Sarcasm Detection Dataset (MUStARD) and introduces a two-phase bimodal data augmentation strategy. The first phase involves generating varied text samples through Back Translation from several secondary languages. The second phase involves the refinement of a FastSpeech 2-based speech synthesis system, tailored specifically for sarcasm to retain sarcastic intonations. Alongside a cloud-based Text-to-Speech (TTS) service, this Fine-tuned FastSpeech 2 system produces corresponding audio for the text augmentations. We also investigate various attention mechanisms for effectively merging text and audio data, finding self-attention to be the most efficient for bimodal integration. Our experiments reveal that this combined augmentation and attention approach achieves a significant F1-score of 81.0% in text-audio modalities, surpassing even models that use three modalities from the MUStARD dataset.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
CHARM: Charge Calibration and Acoustic Rescue for LLM-based Multimodal Sarcasm Detection
Symmetric charged prompts cancel zero-shot LLM sarcasm bias; acoustic late fusion with openSMILE and Omni probes lifts weak backbones up to +0.382 Macro-F1 across English and Chinese.
Reference graph
Works this paper leans on
-
[1]
The role of auditory and visual cues in the interpretation of mandarin ironic speech,
S. Li, A. Chen, Y . Chen, and P. Tang, “The role of auditory and visual cues in the interpretation of mandarin ironic speech,” Journal of Pragmatics, vol. 201, pp. 3–14, Nov. 2022
work page 2022
-
[2]
Tag questions and common ground effects in the perception of verbal irony,
R. J. Kreuz, M. A. Kassler, L. Coppenrath, and B. McLain Allen, “Tag questions and common ground effects in the perception of verbal irony,” Journal of Pragmatics , vol. 31, no. 12, pp. 1685–1700, Nov. 1999
work page 1999
-
[3]
Asymmetries in the use of verbal irony,
R. J. Kreuz and K. E. Link, “Asymmetries in the use of verbal irony,” Journal of Language and Social Psychology, vol. 21, no. 2, pp. 127–143, June 2002
work page 2002
-
[4]
Irony and use-mention distinction,
D. Sperber and D. Wilson, “Irony and use-mention distinction,” in Radical Pragmatics , P. Cole, Ed. New York, NY , USA: Academic Press, 1981, pp. 295–318
work page 1981
-
[5]
On the psycholinguistics of sarcasm,
R. W. Gibbs, “On the psycholinguistics of sarcasm,” Journal of Exper- imental Psychology: General , vol. 115, no. 1, pp. 3–15, March 1986
work page 1986
-
[6]
How to be sarcastic: The reminder theory of verbal irony,
R. J. Kreuz and S. Glucksberg, “How to be sarcastic: The reminder theory of verbal irony,” Journal of Experimental Psychology: General , vol. 118, pp. 347–386, Dec. 1989
work page 1989
-
[7]
Detecting sarcasm in multimodal social platforms,
R. Schifanella, P. de Juan, J. Tetreault, and L. Cao, “Detecting sarcasm in multimodal social platforms,” in Proc. 24th ACM Int. Conf. Multimedia, Amsterdam, The Netherlands, Oct. 2016, pp. 1136–1145
work page 2016
-
[8]
Towards multimodal sarcasm detection (an obviously perfect paper),
S. Castro, D. Hazarika, V . P ´erez-Rosas, R. Zimmermann, R. Mihalcea, and S. Poria, “Towards multimodal sarcasm detection (an obviously perfect paper),” in Proc. 57th Annu. Meeting Assoc. Comput. Linguistics, Florence, Italy, July 2019, pp. 4619–4629
work page 2019
Show all 47 references
-
[9]
Modeling incongruity between modalities for multimodal sarcasm detection,
Y . Wu, Y . Zhao, X. Lu, B. Qin, Y . Wu, J. Sheng, and J. Li, “Modeling incongruity between modalities for multimodal sarcasm detection,”IEEE MultiMedia, vol. 28, no. 2, pp. 86–95, April-June 2021
2021
-
[10]
Multimodal learning using optimal transport for sarcasm and humor detection,
S. Pramanick, A. Roy, and V . M. P. Johns, “Multimodal learning using optimal transport for sarcasm and humor detection,” in 2022 IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), Waikoloa, HI, USA, Jan. 2022, pp. 546–556
2022
-
[11]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. 31st Int. Conf. Neural Inf. Process. Syst., Long Beach, California, USA, June 2017, p. 6000–6010
2017
-
[12]
Improving neural machine translation models with monolingual data,
R. Sennrich, B. Haddow, and A. Birch, “Improving neural machine translation models with monolingual data,” in Proc. 54th Annu. Meeting Assoc. Comput. Linguistics (Volume 1: Long Papers) , Berlin, Germany, Aug. 2016, pp. 86–96
2016
-
[13]
Improving short text classification through global augmentation methods,
V . Marivate and T. Sefara, “Improving short text classification through global augmentation methods,” in Proc. Int. Cross-Domain Conf. Mach. Learn. and Knowl. Extraction (CD-MAKE) , Canterbury, United King- dom, Aug. 2020, pp. 385–399
2020
-
[14]
Audio augmentation for speech recognition,
T. Ko, V . Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition,” in Interspeech 2015 , Dresden, Germany, Sep. 2015, pp. 3586–3589
2015
-
[15]
Training neural speech recognition systems with synthetic speech augmentation,
J. Li, R. T. Gadde, B. Ginsburg, and V . Lavrukhin, “Training neural speech recognition systems with synthetic speech augmentation,” ArXiv, vol. abs/1811.00707, Nov. 2018
2018 arXiv
-
[16]
Mda: Multimodal data aug- mentation framework for boosting performance on sentiment/emotion classification tasks,
N. Xu, W. Mao, P. Wei, and D. Zeng, “Mda: Multimodal data aug- mentation framework for boosting performance on sentiment/emotion classification tasks,” IEEE Intelligent Systems , vol. 36, no. 6, pp. 3–12, Nov.-Dec. 2021
2021
-
[17]
Semantic equivalent adversarial data augmentation for visual question answering,
R. Tang, C. Ma, W. Zhang, Q. Wu, and X. Yang, “Semantic equivalent adversarial data augmentation for visual question answering,” in Proc. 16th Eur. Conf. Comput. Vis. (ECCV) , Glasgow, United Kingdom, Nov. 2020, pp. 437–453
2020
-
[18]
When did you become so smart, oh wise one?! sarcasm explanation in multi-modal multi-party dialogues,
S. Kumar, A. Kulkarni, M. S. Akhtar, and T. Chakraborty, “When did you become so smart, oh wise one?! sarcasm explanation in multi-modal multi-party dialogues,” in Proc. 60th Annu. Meeting Assoc. Comput. Linguistics, Dublin, Ireland, May 2022, pp. 5956–5968
2022
-
[19]
A multimodal corpus for emotion recognition in sarcasm,
A. Ray, S. Mishra, A. Nunna, and P. Bhattacharyya, “A multimodal corpus for emotion recognition in sarcasm,” in Proc. 30th Lang. Resour. and Eval. Conf. , Marseille, France, June 2022, pp. 6992–7003
2022
-
[20]
A multimodal fusion method for sar- casm detection based on late fusion,
N. Ding, S.-w. Tian, and L. Yu, “A multimodal fusion method for sar- casm detection based on late fusion,”Multimedia Tools and Applications, vol. 81, pp. 8597–8616, Feb. 2022
2022
-
[21]
Sarcasm detection using cognitive features of visual data by learning model,
B. N. Hiremath and M. M. Patil, “Sarcasm detection using cognitive features of visual data by learning model,” Expert Systems with Appli- cations, vol. 184, p. 115476, Dec. 2021
2021
-
[22]
Sentiment and emotion help sarcasm? a multi-task learning framework for multi-modal sarcasm, sentiment and emotion analysis,
D. S. Chauhan, D. S. R, A. Ekbal, and P. Bhattacharyya, “Sentiment and emotion help sarcasm? a multi-task learning framework for multi-modal sarcasm, sentiment and emotion analysis,” in Proc. 58th Annu. Meeting Assoc. Comput. Linguistics , Online, July 2020, pp. 4351–4360
2020
-
[23]
Multi-modal sarcasm detection based on contrastive attention mechanism,
X. Zhang, Y . Chen, and G. Li, “Multi-modal sarcasm detection based on contrastive attention mechanism,” in Proc. 10th Natural Lang. Process. and Chin. Comput. , Qingdao, China, Oct 2021, pp. 822–833
2021
-
[24]
Learning multi-task commonness and uniqueness for multi- modal sarcasm detection and sentiment analysis in conversation,
Y . Zhang, Y . Yu, D. Zhao, Z. Li, B. Wang, Y . Hou, P. Tiwari, and J. Qin, “Learning multi-task commonness and uniqueness for multi- modal sarcasm detection and sentiment analysis in conversation,” IEEE Transactions on Artificial Intelligence , pp. 1–13, July 2023
2023
-
[25]
Aggression detection in social media: Using deep neural networks, data augmentation, and pseudo labeling,
S. T. Aroyehun and A. Gelbukh, “Aggression detection in social media: Using deep neural networks, data augmentation, and pseudo labeling,” in Proc. 1st Workshop Trolling, Aggression and Cyberbullying (TRAC) , Santa Fe, New Mexico, USA, Aug. 2018, pp. 90–97
2018
-
[26]
Augmenting data for sarcasm detection with unlabeled conversation context,
H. Lee, Y . Yu, and G. Kim, “Augmenting data for sarcasm detection with unlabeled conversation context,” in Proc. 2nd Workshop Figurative Lang. Process., July 2020, pp. 12–17
2020
-
[27]
Deep cnn-based inductive transfer learning for sarcasm detection in speech,
X. Gao, S. Nayak, and M. Coler, “Deep cnn-based inductive transfer learning for sarcasm detection in speech,” in Interspeech 2022, Incheon, Republic of Korea, Sep. 2022, pp. 2323–2327
2022
-
[28]
Libritts: A corpus derived from librispeech for text-to- speech,
H. Zen, V .-T. Dang, R. A. J. Clark, Y . Zhang, R. J. Weiss, Y . Jia, Z. Chen, and Y . Wu, “Libritts: A corpus derived from librispeech for text-to- speech,” in Interspeech 2019, Graz, Austria, Sep. 2019
2019
-
[29]
Speech recognition with augmented synthesized speech,
A. Rosenberg, Y . Zhang, B. Ramabhadran, Y . Jia, P. Moreno, Y . Wu, and Z. Wu, “Speech recognition with augmented synthesized speech,” in Proc. IEEE Automatic Speech Recognit. and Understanding Workshop (ASRU), Sentosa, Singapore, Dec. 2019, pp. 996–1002
2019
-
[30]
Multimodal continuous emotion recognition with data augmentation using recurrent neural networks,
J. Huang, Y . Li, J. Tao, Z. Lian, M. Niu, and M. Yang, “Multimodal continuous emotion recognition with data augmentation using recurrent neural networks,” in Proc. Audio/Vis. Emotion Challenge and Workshop, Seoul, Republic of Korea, Oct. 2018, pp. 57–64
2018
-
[31]
Mixgen: A new multi-modal data augmentation,
X. Hao, Y . Zhu, S. Appalaraju, A. Zhang, W. Zhang, B. Li, and M. Li, “Mixgen: A new multi-modal data augmentation,” in Proc. IEEE/CVF Winter Conf. Appl. Comput. Vis. Workshops (WACVW) , Waikoloa, Hawaii, Feb. 2023, pp. 379–389
2023
-
[32]
Bert: Pre-training of deep bidirectional transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proc. North American Chapter Assoc. Comput. Linguistics (NAACL) , Hyatt Regency, Minneapolis, USA, June 2019, pp. 4171–4186
2019
-
[33]
Cnn architectures for large-scale audio classifica- tion,
S. Hershey and et al, “Cnn architectures for large-scale audio classifica- tion,” in Proc. 42nd IEEE Int. Conf. Acoust., Speech and Signal Process. (ICASSP), New Orleans, USA, March 2017, pp. 131–135
2017
-
[34]
Inves- tigating on incorporating pretrained and learnable speaker representa- tions for multi-speaker multi-style text-to-speech,
C.-M. Chien, J.-H. Lin, C.-y. Huang, P.-c. Hsu, and H.-y. Lee, “Inves- tigating on incorporating pretrained and learnable speaker representa- tions for multi-speaker multi-style text-to-speech,” in Proc. IEEE Int. xiii Conf. Acoust., Speech and Signal Process. (ICASSP) , Toron...
2021
-
[35]
Adam: A method for stochastic optimization,
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. 3rd Int. Conf. Learn. Representations (ICLR) , San Diego, CA, USA, Dec. 2011
2011
-
[36]
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Proc. 34th Int. Conf. Neural Inf. Process. Systems , Vancouver, BC, Canada, Oct. 2020
2020
-
[37]
A quantum probability driven frame- work for joint multi-modal sarcasm, sentiment and emotion analysis,
Y . Liu, Y . Zhang, and D. Song, “A quantum probability driven frame- work for joint multi-modal sarcasm, sentiment and emotion analysis,” IEEE Transactions on Affective Computing , pp. 1–15, 2023
2023
-
[38]
A multitask learning model for multimodal sarcasm, sentiment and emotion recognition in conversations,
Y . Zhang, J. Wang, Y . Liu, L. Rong, Q. Zheng, D. Song, P. Tiwari, and J. Qin, “A multitask learning model for multimodal sarcasm, sentiment and emotion recognition in conversations,” Information Fusion, vol. 93, pp. 282–301, May 2023
2023
-
[39]
Support-vector networks,
C. Cortes and V . Vapnik, “Support-vector networks,” Machine Learning, vol. 20, no. 3, pp. 273–297, Sep. 1995
1995
-
[40]
An emoji-aware multitask framework for multimodal sarcasm detec- tion,
D. S. Chauhan, G. V . Singh, A. Arora, A. Ekbal, and P. Bhattacharyya, “An emoji-aware multitask framework for multimodal sarcasm detec- tion,” Knowledge-Based Systems, vol. 257, Dec. 2022
2022
-
[41]
Dropout: A simple way to prevent neural networks from overfit- ting,
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhut- dinov, “Dropout: A simple way to prevent neural networks from overfit- ting,” Journal of Machine Learning Research , vol. 15, no. 1, pp. 1929– 1958, Jan. 2014
1929
-
[42]
Rectified linear units improve restricted boltzmann machines,
V . Nair and G. E. Hinton, “Rectified linear units improve restricted boltzmann machines,” in Proc. 27th Int. Conf. Mach. Learn. , Haifa, Israel, June 2010, pp. 807–814
2010
-
[43]
Glove: Global vectors for word representation,
J. Pennington, R. Socher, and C. Manning, “Glove: Global vectors for word representation,” in Proc. 2014 Conf. Empirical Methods Natural Lang. Process. (EMNLP) , Doha, Qatar, Oct 2014, pp. 1532–1543
2014
-
[44]
Generative emotional ai for speech emotion recognition: The case for synthetic emotional speech augmen- tation,
S. Latif, A. Shahid, and J. Qadir, “Generative emotional ai for speech emotion recognition: The case for synthetic emotional speech augmen- tation,” Applied Acoustics, vol. 210, July 2023
2023
-
[45]
Multimodal markers of irony and sarcasm,
S. Attardo, J. Eisterhold, J. Hay, and I. Poggi, “Multimodal markers of irony and sarcasm,” Humor - International Journal of Humor Research , vol. 16, no. 2, Jan. 2003
2003
-
[46]
Exploring the role body in communicating ironic stance,
C. de Vries, B. Oben, and G. Br ˆone, “Exploring the role body in communicating ironic stance,” Languages and Modalities , vol. 1, pp. 65–80, Oct. 2021. Xiyuan Gao received the Master’s degree in Speech and Language Processing from Konstanz University, Konstanz, Germany. She i...
2021
-
[2010]
In 2012, he returned to academia
After a postdoctoral appointment at his alma mater, he joined an AI start-up focusing on acoustic sensors, where he served as Head of the Cognitive Systems Unit. In 2012, he returned to academia. Dr. Coler is currently an Associate Professor of Speech Technology at the Faculty...
2012
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.