REVIEW 4 major objections 6 minor 77 references
Looking at the user’s face lets a speech synthesizer produce more natural, emotionally fitting replies in conversation.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 15:04 UTC pith:22ZYC5VF
load-bearing objection Solid multimodal CSS systems paper with a real dataset and a compact AU tokenizer; the face-causality claim is under-supported, but the engineering package still deserves referee time. the 4 major comments →
Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
FacialTalker shows that encoding each user face frame as one Action-Unit-supervised discrete token, modeling it jointly with text and speech inside an LLM, and post-training with dual preference constraints on face and speech sequences, yields conversational speech that is more natural, expressive, and context-aligned than strong baselines that lack this facial pathway—or that use coarser visual encoders.
What carries the argument
AUTokenizer: a single-codebook Finite Scalar Quantization tokenizer that maps each facial frame to one discrete token under supervision from combinations of facial Action Units, plus DualDPO, which extends direct preference optimization to rank both visual and speech token sequences together.
Load-bearing premise
The load-bearing premise is that Action Unit labels from small micro-expression corpora, compressed into one token per frame, carry enough spontaneous conversational affect to improve target emotion and prosody—not merely that more data or speech-side training would have produced the same gains.
What would settle it
Retrain FacialTalker on the same large dialogue data but replace AUTokenizer with a scrambled or constant face token (or drop DualDPO’s visual preference term) and check whether emotion accuracy, prosody distance, and human MOS_E on MultiDialog, AvaMERG, and VSDD-1K collapse back to non-visual or CLIP-based baselines.
If this is right
- Empathetic CSS systems can treat user face video as a first-class context stream without multi-token image codecs.
- A single AU-supervised face token per frame is a usable interface between facial analysis and autoregressive speech LLMs.
- Joint preference optimization on visual and speech tokens is a transferable post-training recipe for multimodal dialogue agents.
- The automated VSDD-1K pipeline supplies open, large-scale face-aligned dialogue data for other speech and affect tasks.
- Agents in homes, cockpits, and eldercare can condition replies on how the user looks, not only on what they say.
Where Pith is reading between the lines
- If one AU token suffices here, the same compression may unlock face-conditioned spoken dialogue on smaller on-device LLMs where multi-token vision is too costly.
- Failures will likely concentrate on AU combinations rare in CASME II/DISFA but common in spontaneous talk (e.g., polite smiles vs genuine joy), suggesting a need for in-the-wild AU labels.
- DualDPO’s visual rejected samples are model-generated face tokens; the same idea could regularize other sparse nonverbal streams such as gaze or posture tokens.
- Releasing 1K hours of real face-to-face talk may matter as much as the model for follow-on work on turn-taking and micro-expression-aware dialogue.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents FacialTalker, an LLM-based (Qwen2.5-0.5B, initialized from CosyVoice2) conversational speech synthesis system that conditions on user facial expressions in addition to text/speech history. Three main components are claimed: (i) AUTokenizer, a ConvNeXt-Tiny + FSQ visual tokenizer that compresses each face frame into a single discrete token supervised by Action Unit combination labels from CASME II/DISFA; (ii) DualDPO, a post-training stage applying DPO preference pairs to both speech and facial token sequences (rejected samples drawn from the Stage-3 model's own outputs); and (iii) VSDD-1K, a 1,033-hour automatically constructed video-speech dialogue dataset from interviews/podcasts with >85% valid-face frames. Evaluation on MultiDialog, AvaMERG, and VSDD-1K (Tables 4–5) shows consistent gains over non-visual CSS baselines and visual baselines (Empatheia, EmpathyEar, UniTalker) in SIM, PDTW, ACC_E, and MOS with 95% CIs from 50 raters, supported by ablations (FT-base, FT-CLIP, AUTokenizer-VQ) and AU F1 benchmarks (Table 3).
Significance. If the central claim holds, this is a useful contribution: it is, to my knowledge, the first LLM-based CSS system to incorporate frame-level facial affect at this scale, and VSDD-1K — a 1K-hour open-source synchronized face-speech dialogue corpus with a documented, reproducible construction pipeline (§5.1) — is independently valuable to the community. The experimental apparatus is substantial for the venue: three datasets, objective metrics external to the training losses (WavLM/CAM++ SIM, Emotion2vec ACC_E, PDTW), 50-rater MOS with confidence intervals, ablations of both proposed components, LOSO cross-validation on AU corpora, and a promised code/dataset release. The single-token-per-frame FSQ design is a clean, cheap interface to LLM sequence modeling and the DualDPO extension is a natural idea. The main limitation on significance is attribution: the experiments as designed cannot yet isolate the facial modality's causal contribution, and the head-to-head with the strongest baseline (UniTalker) is confounded.
major comments (4)
- [§7.4 Ablation Study] §7.4 / Tables 4–5: there is no ablation that removes, masks, or shuffles the facial input at inference or training, so the paper never demonstrates that the facial modality is causally responsible for the gains over text/speech-only context. The ablations vary the visual encoder (FT-CLIP) and remove DualDPO (FT-base), but every condition still receives aligned faces. A face-shuffle or <IGNORE>-mask condition is essentially free — the <IGNORE> token mechanism already exists (§5.1.4) — and is load-bearing for the abstract's claim that facial cues make the output 'more natural, expressive, and better aligned with the conversational context.' If performance is nearly unchanged under shuffled faces (a plausible outcome given how much context the speech/text history carries in CSS), the framing of the paper changes materially. This experiment must be added.
- [§7.3] The sentence 'FT-base and UniTalker mainly differ in their visual tokenizer design' is not supported. FacialTalker is Qwen2.5-0.5B initialized from CosyVoice2 (170k h pretraining, §4.3.1) trained on ~1,724 h with a four-stage curriculum; UniTalker is a different architecture, initialization, and training corpus. The SIM 0.90→0.92 and PDTW 42.01→39.29 gaps on MultiDialog can plausibly come from backbone, initialization, or data scale rather than AU tokens vs. 128 landmarks. Either provide a matched-backbone comparison (e.g., swap landmark tokens into the same LLM pipeline) or explicitly downgrade this to a systems comparison and remove the tokenizer-attribution language.
- [§4.1.3 / §6.1 / Table 3] Two linked issues. (1) AUTokenizer's AU-supervision comes from CASME II (247 samples, 26 subjects) and DISFA (27 subjects) — small, posed/micro-expression corpora with non-overlapping AU label sets (§6.1) — yet the tokenizer is deployed on spontaneous podcast/interview faces in VSDD-1K. No measurement of token quality exists on the target domain (no AU labels there, no reconstruction or probe metric on VSDD-1K faces). The transfer assumption is asserted, not tested. (2) Table 3 shows AUTokenizer is second-tier as an AU classifier: 0.65 avg F1 on DISFA vs. VL-FAU 0.66, and 0.75 on CASME II vs. AULLM 0.81 and SSSNet-LED 0.79. That is acceptable for a tokenizer whose real job is compression, but the paper should then provide direct evidence that the single FSQ token (levels [8,8,8,8,8], i.e., ≤32,768 codes) retains affectively relevant information — e.g., codebook utilization, emotion probe
- [§4.3.1 DualDPO] The DualDPO gains, while consistent, are modest (e.g., MultiDialog ACC_E 0.76→0.79, MOS_N 4.16→4.22 over FT-base), and the rejected samples are the Stage-3 model's own generations. This self-preference construction is standard practice, but it carries a known risk of reward hacking toward the model's own distribution, and the paper gives no analysis of what the visual-side DPO actually changes (e.g., do chosen-vs-rejected facial token sequences differ in ways correlated with downstream emotion accuracy?). A brief analysis of preference-pair statistics and a sentence on why self-generated rejected samples are appropriate here would strengthen §4.3.1; if the visual DPO contributes little independently of the speech DPO, that should be reported.
minor comments (6)
- [Tables 4–5] Objective metrics in Tables 4–5 report no variance or CIs (only MOS has them). Given that some margins are small (SIM 0.91 vs 0.92 on VSDD-1K; FT-base vs full PDTW 41.52 vs 40.10), please report seed variance or a paired significance test over the 200-sample evaluation sets.
- [§6.3] The MOS raters are described as 'participants who speak English as a second language.' For MOS_N (naturalness) judgments of English speech, please justify this choice or report a subset with native-speaker raters; naturalness ratings are known to be rater-population-sensitive.
- [§7.5 / Fig. 5] Fig. 5's attention-visualization methodology is unspecified (which layer/heads are visualized for AUTokenizer? how is CLIP's 'facial expression' prompt attention computed and made comparable to a single-token FSQ model?). As presented the comparison is suggestive but not controlled.
- [§4.1.3] The number N of learnable AU queries (§4.1.3) is never stated, nor are training hyperparameters, inference cost, or latency of the full pipeline. Given that efficiency ('significantly reducing training and inference costs') is a stated motivation, a parameter/FLOPs/latency table versus UniTalker and the CLIP variant would be appropriate.
- [§3 / Table 1 / Eq. (2)] Typos/notation: §3 'should be not only natural' is missing a 'be'; Table 1's 'Modal' column renders as '( ,a,t)' with a missing modality symbol; Eq. (2) would benefit from a sentence of explanation; the alignment between 25 FPS face tokens and the speech token rate (and the <IGNORE> padding ratio in practice) should be stated explicitly since it affects context length.
- [§6.1] VSDD-1K test evaluation is an 8:1:1 random split of the same distribution used for training (§6.1). A held-out-source or held-out-speaker evaluation (VSDD-1K has >1498 speakers) would substantially strengthen the generalization claims and is feasible with the data already collected.
Circularity Check
No derivation-chain circularity: claims are empirical system results evaluated on external metrics, not predictions forced by their inputs.
full rationale
FacialTalker is an engineering/ML systems paper. Its load-bearing claims are comparative empirical results (Tables 3–5: AU F1, SIM, PDTW, ACC_E, MOS) on held-out dialogue data against external baselines and human raters. AUTokenizer is supervised by independent AU combination labels from CASME II/DISFA and scored with F1; speech quality uses third-party metrics (WavLM, CAM++, Emotion2vec, DNSMOS, UTMOS) and human MOS—none of which equal the training losses by construction. DualDPO (§4.3.1) uses Stage-3 model outputs as rejected pairs and ground-truth tokens as chosen pairs; that is standard preference optimization, not a fitted parameter renamed as a prediction, and it does not make the reported test metrics tautological. Self-citations to the authors’ prior CSS stack (GPT-Talker, ECSS, UniTalker, etc.) appear as related work and baselines, not as uniqueness theorems or ansatzes that force the present result. There is no self-definitional equation, no uniqueness import, and no renaming of a known closed-form result. Experimental gaps (missing face-shuffle/removal ablations, backbone confounds) are correctness/attribution issues, not circularity of the derivation chain. Score 0; steps empty.
Axiom & Free-Parameter Ledger
free parameters (5)
- FSQ levels [8,8,8,8,8] (single facial token codebook geometry) =
[8,8,8,8,8]
- Number of learnable AU queries N and 128-d fused AU feature size =
N AU queries; 128-d concat (paper)
- DualDPO preference-pair construction (Stage-3 self-samples as rejected)
- Pipeline thresholds (VAD pause >5s; SNR-related cutoff 'below 4'; 25 FPS / 16 kHz) =
pause>5s; threshold 4; 25fps/16kHz mono
- LLM backbone scale and init (Qwen2.5-0.5B from CosyVoice2 170k-h pretrain) =
Qwen2.5-0.5B; CosyVoice2 init
axioms (6)
- domain assumption Combinations of facial Action Units are a sufficient supervisory signal for conversational facial affect relevant to empathetic speech style.
- ad hoc to paper One discrete token per frame is enough visual bandwidth for LLM autoregressive multimodal context modeling in CSS.
- domain assumption User-only face streams (agent face omitted) match real deployment and still supply the affect needed for agent speech.
- domain assumption Automatically collected interview/podcast dialogues with ASD-filtered faces are a valid large-scale proxy for natural multimodal conversation.
- domain assumption Standard next-token CE plus DPO-style preference optimization improves joint visual–speech contextual understanding.
- domain assumption BPE text tokens, SenseVoice+FSQ speech tokens, and emotion category tokens form a compatible discrete interface with face tokens for a single LLM.
invented entities (4)
-
AUTokenizer
independent evidence
-
DualDPO
no independent evidence
-
FacialTalker
no independent evidence
-
VSDD-1K
independent evidence
read the original abstract
Conversational Speech Synthesis is a fundamental component of human-computer interaction, aiming to generate contextually appropriate, expressive, and empathetic speech. However, facial expressions encode subtle and rich affective cues that are crucial for empathetic speech interaction, whereas existing approaches often overlook this important modality. In addition, the lack of large-scale natural conversational datasets with both speech and visual modalities also limits the development of visual affect understanding in conversational settings.To address these limitations, we propose FacialTalker, a facial-expression-aware CSS framework built upon a large language model backbone. To efficiently encode facial expressions, we propose AUTokenizer, a single-codebook visual tokenizer that discretizes each frame-level facial expression into a compact token, trained with supervision from combinations of facial Action Units. We further introduce a dual direct preference optimization (DualDPO) strategy, which extends the DPO by jointly imposing preference constraints on both visual and speech token sequences, to enhance the model's understanding of facial expressions and speech semantics in multimodal conversational contexts. Moreover, we construct VSDD-1K, a large-scale multimodal dialogue dataset collected through a fully automated pipeline from real-world Internet conversations, comprising over 1,033 hours of synchronized speaker videos and speech, with more than 85\% of frames containing valid faces. Extensive objective and subjective experiments demonstrate that FacialTalker consistently outperforms strong baselines in facial-expression perception and speech synthesis quality, generating speech that is more natural, expressive, and better aligned with the conversational context. The results also validate the effectiveness of our training strategy and dataset construction pipeline.
Figures
Reference graph
Works this paper leans on
-
[1]
Inclusion AI, Bowen Ma, Cheng Zou, Canxiang Yan, Chunxiang Jin, Chunjie Shen, Chenyu Lian, Dandan Zheng, Fudong Wang, Furong Xu, et al. 2025. Ming-flash- omni: A sparse, unified architecture for multimodal perception and generation. arXiv preprint arXiv:2510.24821(2025)
arXiv 2025
-
[2]
Keyu An, Qian Chen, Chong Deng, Zhihao Du, Changfeng Gao, Zhifu Gao, Yue Gu, Ting He, Hangrui Hu, Kai Hu, et al. 2024. Funaudiollm: Voice understanding and generation foundation models for natural interaction between humans and llms.arXiv preprint arXiv:2407.04051(2024)
Pith/arXiv arXiv 2024
-
[3]
2022.Face-to-face dialogue: Theory, research, and applica- tions
Janet Beavin Bavelas. 2022.Face-to-face dialogue: Theory, research, and applica- tions. Oxford University Press
2022
-
[4]
Fadi Boutros, Meiling Fang, Marcel Klemt, Biying Fu, and Naser Damer. 2023. CR- FIQA: face image quality assessment by learning sample relative classifiability. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 5836–5845
2023
-
[5]
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan. 2008. IEMOCAP: Interactive emotional dyadic motion capture database.Language resources and evaluation42, 4 (2008), 335–359
2008
-
[6]
Rajdeep Chatterjee, Saptarshi Mazumdar, R Simon Sherratt, Rohit Halder, Tanmoy Maitra, and Debasis Giri. 2021. Real-time speech emotion analysis for smart home assistants.IEEE Transactions on Consumer Electronics67, 1 (2021), 68–76
2021
-
[7]
Sanyuan Chen, Shujie Liu, Long Zhou, Yanqing Liu, Xu Tan, Jinyu Li, Sheng Zhao, Yao Qian, and Furu Wei. 2024. Vall-e 2: Neural codec language models are human parity zero-shot text to speech synthesizers.arXiv preprint arXiv:2406.05370 (2024)
Pith/arXiv arXiv 2024
-
[8]
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al . 2022. Wavlm: Large-scale self-supervised pre-training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing16, 6 (2022), 1505–1518
2022
-
[9]
Yirong Chen, Weiquan Fan, Xiaofen Xing, Jianxin Pang, Minlie Huang, Wenjing Han, Qianfeng Tie, and Xiangmin Xu. 2022. CPED: A large-scale Chinese per- sonalized and emotional dialogue dataset for conversational AI.arXiv preprint arXiv:2205.14727(2022)
Pith/arXiv arXiv 2022
-
[10]
Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, and Neil Zeghidour. 2024. Moshi: a speech-text foundation model for real-time dialogue.arXiv preprint arXiv:2410.00037(2024)
Pith/arXiv arXiv 2024
-
[11]
Zhihao Du, Yuxuan Wang, Qian Chen, Xian Shi, Xiang Lv, Tianyu Zhao, Zhifu Gao, Yexin Yang, Changfeng Gao, Hui Wang, et al . 2024. Cosyvoice 2: Scal- able streaming speech synthesis with large language models.arXiv preprint arXiv:2412.10117(2024)
Pith/arXiv arXiv 2024
-
[12]
Patrick Esser, Robin Rombach, and Bjorn Ommer. 2021. Taming transformers for high-resolution image synthesis. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 12873–12883
2021
-
[13]
Hao Fei, Han Zhang, Bin Wang, Lizi Liao, Qian Liu, and Erik Cambria. 2024. Empathyear: An open-source avatar multimodal empathetic chatbot. InProceed- ings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations). 61–71
2024
-
[14]
Xuri Ge, Junchen Fu, Fuhai Chen, Shan An, Nicu Sebe, and Joemon M Jose
-
[15]
Haohan Guo, Shaofei Zhang, Frank K Soong, Lei He, and Lei Xie. 2021. Conversa- tional end-to-end tts for voice agents. In2021 IEEE Spoken Language Technology Workshop (SLT). IEEE, 403–409
2021
-
[16]
Guohong Hu, Xing Lan, Hanyu Jiang, Jiayi Lyu, and Jian Xue. 2024. Towards unified facial action unit recognition framework by large language models.arXiv preprint arXiv:2409.08444(2024)
Pith/arXiv arXiv 2024
-
[17]
Hangrui Hu, Xinfa Zhu, Ting He, Dake Guo, Bin Zhang, Xiong Wang, Zhifang Guo, Ziyue Jiang, Hongkun Hao, Zishan Guo, et al. 2026. Qwen3-TTS Technical Report.arXiv preprint arXiv:2601.15621(2026)
Pith/arXiv arXiv 2026
-
[18]
Yifan Hu, Rui Liu, Guanglai Gao, and Haizhou Li. 2024. FCTalker: Fine and Coarse Grained Context Modeling for Expressive Conversational Speech Synthesis. In 14th IEEE International Symposium on Chinese Spoken Language Processing, ISCSLP 2024, Beijing, China, November 7-10, 2024, Yanmin Qian, Qin Jin, Zhijian Ou, Zhenhua Ling, Zhiyong Wu, Ya Li, Lei Xie, a...
arXiv 2024
-
[19]
Yifan Hu, Rui Liu, Yi Ren, Xiang Yin, and Haizhou Li. 2025. UniTalker: Conver- sational Speech-Visual Synthesis. InProceedings of the 33rd ACM International Conference on Multimedia. 10248–10257
2025
-
[20]
Byungho Jo, Donghyeon Cho, In Kyu Park, and Sungeun Hong. 2023. IFQA: Interpretable face quality assessment. InProceedings of the IEEE/CVF winter conference on applications of computer vision. 3444–3453
2023
-
[21]
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.Advances in neural information processing systems33 (2020), 17022–17033
2020
-
[22]
Keon Lee, Kyumin Park, and Daeyoung Kim. 2023. Dailytalk: Spoken dialogue dataset for conversational text-to-speech. InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5
2023
-
[23]
Jidong Leng and Qiang Yan. 2025. Eijl: Popularity prediction of social media advertisements based on multimodal emotional interaction and joint learning. Data Intelligence7, 4 (2025), 1129–1146
2025
-
[24]
Jingbei Li, Yi Meng, Chenyi Li, Zhiyong Wu, Helen Meng, Chao Weng, and Dan Su. 2022. Enhancing speaking styles in conversational text-to-speech synthesis with graph-based multi-modal context modeling. InICASSP 2022-2022 IEEE In- ternational Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 7917–7921
2022
-
[25]
Jingbei Li, Yi Meng, Xixin Wu, Zhiyong Wu, Jia Jia, Helen Meng, Qiao Tian, Yup- ing Wang, and Yuxuan Wang. 2022. Inferring speaking styles from multi-modal conversational context by multi-scale relational graph convolutional networks. In Proceedings of the 30th ACM International Conference on Multimedia. 5811–5820
2022
-
[26]
Yifan Li, Anh Dao, Wentao Bao, Zhen Tan, Tianlong Chen, Huan Liu, and Yu Kong
-
[27]
Yante Li, Xiaohua Huang, and Guoying Zhao. 2021. Micro-expression action unit detection with spatial and channel attention.Neurocomputing436 (2021), 221–231
2021
-
[28]
InEuropean Conference on Computer Vision
Facial affective behavior analysis with instruction tuning. InEuropean Conference on Computer Vision. Springer, 165–186
-
[29]
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. Dailydialog: A manually labelled multi-turn dialogue dataset. InProceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 986–995
2017
-
[30]
Yante Li, Wei Peng, and Guoying Zhao. 2021. Micro-expression action unit de- tection with dual-view attentive similarity-preserving knowledge distillation. In 2021 16th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2021). IEEE, 01–08
2021
-
[31]
Guan-Ting Lin, Prashanth Gurunath Shivakumar, Ankur Gandhe, Chao- Han Huck Yang, Yile Gu, Shalini Ghosh, Andreas Stolcke, Hung-yi Lee, and Ivan Bulyko. 2024. Paralinguistics-enhanced large language modeling of spoken dialogue. InICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 10316–10320
2024
-
[32]
Junhua Liao, Haihan Duan, Kanghui Feng, Wanbing Zhao, Yanbing Yang, Liangyin Chen, and Yanru Chen. 2025. Lr-asd: Lightweight and robust network for active speaker detection.International Journal of Computer Vision133, 7 (2025), 4749– 4769
2025
-
[33]
Rui Liu, Yifan Hu, Yi Ren, Xiang Yin, and Haizhou Li. 2024. Emotion rendering for conversational speech synthesis with heterogeneous graph-based context modeling. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 18698–18706
2024
-
[34]
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le
-
[35]
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. 2022. A convnet for the 2020s. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 11976–11986
2022
-
[36]
Zhishu Liu, Kaishen Yuan, Bo Zhao, Yong Xu, and Zitong Yu. 2025. Au-llm: Micro-expression action unit detection via enhanced llm-based feature fusion. In Chinese Conference on Biometric Recognition. Springer, 355–365
2025
-
[37]
Rui Liu, Yifan Hu, Yi Ren, Xiang Yin, and Haizhou Li. 2024. Generative expressive conversational speech synthesis. InProceedings of the 32nd ACM International Conference on Multimedia. 4187–4196
2024
-
[38]
Ziyang Ma, Zhisheng Zheng, Jiaxin Ye, Jinchao Li, Zhifu Gao, Shiliang Zhang, and Xie Chen. 2024. emotion2vec: Self-supervised pre-training for speech emotion MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil Yifan Hu, Shuwei He, Rui Liu, and Haizhou Li. representation. InFindings of the Association for Computational Linguistics: ACL
2024
-
[39]
Brais Martinez, Michel F Valstar, Bihan Jiang, and Maja Pantic. 2017. Automatic analysis of facial actions: A survey.IEEE transactions on affective computing10, 3 (2017), 325–347
2017
-
[40]
Cheng Luo, Siyang Song, Weicheng Xie, Linlin Shen, and Hatice Gunes. 2022. Learning multi-dimensional edge feature-based au relation graph for facial action unit recognition.arXiv preprint arXiv:2205.01782(2022)
Pith/arXiv arXiv 2022
-
[41]
Fabian Mentzer, David Minnen, Eirikur Agustsson, and Michael Tschannen. 2023. Finite scalar quantization: Vq-vae made simple.arXiv preprint arXiv:2309.15505 (2023)
Pith/arXiv arXiv 2023
-
[42]
Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. 2016. V-net: Fully convolutional neural networks for volumetric medical image segmentation. In 2016 fourth international conference on 3D vision (3DV). Ieee, 565–571
2016
-
[43]
S Mohammad Mavadati, Mohammad H Mahoor, Kevin Bartlett, Philip Trinh, and Jeffrey F Cohn. 2013. Disfa: A spontaneous facial action intensity database.IEEE Transactions on Affective Computing4, 2 (2013), 151–160
2013
-
[44]
Se Park, Chae Kim, Hyeongseop Rha, Minsu Kim, Joanna Hong, Jeonghun Yeo, and Yong Ro. 2024. Let’s go real talk: Spoken dialogue model for face-to-face conversation. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 16334–16348
2024
-
[45]
Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea. 2019. Meld: A multimodal multi-party dataset for emotion recognition in conversations. InProceedings of the 57th annual meeting of the association for computational linguistics. 527–536
2019
-
[46]
Jinhui Pang, Xinyun Yang, Xiaoyao Qiu, Zixuan Wang, and Taisheng Huang
-
[47]
MMAF: Masked Multi-modal Attention Fusion to Reduce Bias of Visual Features for Named Entity Recognition.Data Intelligence6, 4 (2024), 1114–1133
2024
-
[48]
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2020. Fastspeech 2: Fast and high-quality end-to-end text to speech.arXiv preprint arXiv:2006.04558(2020)
Pith/arXiv arXiv 2020
-
[49]
Tal Ridnik, Emanuel Ben-Baruch, Nadav Zamir, Asaf Noy, Itamar Friedman, Matan Protter, and Lihi Zelnik-Manor. 2021. Asymmetric loss for multi-label classification. InProceedings of the IEEE/CVF international conference on computer vision. 82–91
2021
-
[50]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. InInternational conference on machine learning. PmLR, 8748–8763
2021
-
[51]
Chandan KA Reddy, Vishak Gopal, and Ross Cutler. 2021. DNSMOS: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors. InICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 6493–6497
2021
-
[52]
Yuki Saito, Yuto Nishimura, Shinnosuke Takamichi, Kentaro Tachibana, and Hiroshi Saruwatari. 2022. STUDIES: Corpus of Japanese empathetic dialogue speech towards friendly voice agent.arXiv preprint arXiv:2203.14757(2022)
Pith/arXiv arXiv 2022
-
[53]
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. InProceedings of the 54th annual meeting of the association for computational linguistics (volume 1: long papers). 1715–1725
2016
-
[54]
Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, and Hiroshi Saruwatari. 2022. Utmos: Utokyo-sarulab system for voicemos challenge 2022.arXiv preprint arXiv:2204.02152(2022)
Pith/arXiv arXiv 2022
-
[55]
Yuki Saito, Eiji Iimori, Shinnosuke Takamichi, Kentaro Tachibana, and Hiroshi Saruwatari. 2023. CALLS: Japanese empathetic dialogue speech corpus of complaint handling and attentive listening in customer center.arXiv preprint arXiv:2305.13713(2023)
Pith/arXiv arXiv 2023
-
[56]
Junke Wang, Yi Jiang, Zehuan Yuan, Binyue Peng, Zuxuan Wu, and Yu-Gang Jiang. 2024. Omnitokenizer: A joint image-video tokenizer for visual generation. Advances in Neural Information Processing Systems37 (2024), 28281–28295
2024
-
[57]
Ning Wang, Frank Broz, Alessandro Di Nuovo, Tony Belpaeme, and Angelo Cangelosi. 2016. A user-centric design of service robots speech interface for the elderly. InRecent Advances in Nonlinear Speech Processing. Springer, 275–283
2016
-
[58]
Meituan LongCat Team, Bairui Wang, Bin Xiao, Bo Zhang, Bolin Rong, Borun Chen, Chang Wan, Chao Zhang, Chen Huang, Chen Chen, et al. 2025. Longcat- flash-omni technical report.arXiv preprint arXiv:2511.00279(2025)
arXiv 2025
-
[59]
Tuomas Varanka, Wei Peng, and Guoying Zhao. 2023. Learnable eulerian dy- namics for micro-expression action unit detection. InScandinavian Conference on Image Analysis. Springer, 385–400
2023
-
[60]
Zhifei Xie and Changqiao Wu. 2024. Mini-omni2: Towards open-source gpt- 4o with vision, speech and duplex capabilities.arXiv preprint arXiv:2410.11190 (2024)
Pith/arXiv arXiv 2024
-
[61]
Jin Xu, Zhifang Guo, Hangrui Hu, Yunfei Chu, Xiong Wang, Jinzheng He, Yuxuan Wang, Xian Shi, Ting He, Xinfa Zhu, et al. 2025. Qwen3-omni technical report. arXiv preprint arXiv:2509.17765(2025)
Pith/arXiv arXiv 2025
-
[62]
Boyong Wu, Chao Yan, Chen Hu, Cheng Yi, Chengli Feng, Fei Tian, Feiyu Shen, Gang Yu, Haoyang Zhang, Jingbei Li, et al. 2025. Step-audio 2 technical report. arXiv preprint arXiv:2507.16632(2025)
Pith/arXiv arXiv 2025
-
[63]
Lei Wu, Jiyong Xue, Wenbo Li, Kan Wang, Xiang Zhang, and Gang Guo. 2022. Toward decreasing the driving risk: speech-based driver’s anger regulation in smart cockpit.IEEE Journal of Radio Frequency Identification6 (2022), 764–768
2022
-
[64]
Wen-Jing Yan, Xiaobai Li, Su-Jing Wang, Guoying Zhao, Yong-Jin Liu, Yu-Hsin Chen, and Xiaolan Fu. 2014. CASME II: An improved spontaneous micro- expression database and the baseline evaluation.PloS one9, 1 (2014), e86041
2014
-
[65]
Zehui Yang, Yifan Chen, Lei Luo, Runyan Yang, Lingxuan Ye, Gaofeng Cheng, Ji Xu, Yaohui Jin, Qingqing Zhang, Pengyuan Zhang, et al. 2022. Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational (RAMC) Speech Dataset. InProc. Interspeech 2022. 1736–1740
2022
-
[66]
Hongfei Xue, Yuhao Liang, Bingshen Mu, Shiliang Zhang, Mengzhe Chen, Qian Chen, and Lei Xie. 2024. E-chat: Emotion-sensitive spoken dialogue system with large language models. In2024 IEEE 14th International Symposium on Chinese Spoken Language Processing (ISCSLP). IEEE, 586–590
2024
-
[67]
Jinlong Xue, Yayue Deng, Fengping Wang, Ya Li, Yingming Gao, Jianhua Tao, Jianqing Sun, and Jiaen Liang. 2023. M 2-ctts: End-to-end multi-scale multi-modal conversational text-to-speech synthesis. InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5
2023
-
[68]
Fan Zhang and Lin Chai. 2024. A review of research on micro-expression recog- nition algorithms based on deep learning.Neural Computing and Applications36, 29 (2024), 17787–17828
2024
-
[69]
Han Zhang, Zixiang Meng, Meng Luo, Hong Han, Lizi Liao, Erik Cambria, and Hao Fei. 2025. Towards multimodal empathetic response generation: A rich text-speech-vision avatar-based benchmark. InProceedings of the ACM on Web Conference 2025. 2872–2881
2025
-
[70]
Rohola Zandie, Mohammad H Mahoor, Julia Madsen, and Eshrat S Emamian
-
[71]
Siyi Zhou, Yiquan Zhou, Yi He, Xun Zhou, Jinchao Wang, Wei Deng, and Jingchen Shu. 2026. Indextts2: A breakthrough in emotionally expressive and duration- controlled auto-regressive zero-shot text-to-speech. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 35139–35148
2026
-
[72]
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. 2023. Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities. InFindings of the Association for Computational Linguistics: EMNLP 2023. 15757–15773
2023
-
[75]
Jinming Zhao, Tenggan Zhang, Jingwen Hu, Yuchen Liu, Qin Jin, Xinchao Wang, and Haizhou Li. 2022. M3ED: Multi-modal multi-scene multi-label emotional dialogue database. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 5699–5710
2022
-
[77]
Yixuan Zhou, Guoyang Zeng, Xin Liu, Xiang Li, Renjie Yu, Ziyang Wang, Runchuan Ye, Weiyue Sun, Jiancheng Gui, Kehan Li, et al . 2025. Voxcpm: Tokenizer-free TTS for context-aware speech generation and true-to-life voice cloning.arXiv preprint arXiv:2509.24650(2025)
arXiv 2025
-
[2021]
RyanSpeech: A Corpus for Conversational Text-to-Speech Synthesis. In Proc. Interspeech 2021. 2751–2755
2021
-
[2022]
Flow matching for generative modeling.arXiv preprint arXiv:2210.02747 (2022)
Pith/arXiv arXiv 2022
-
[2024]
InProceedings of the 32nd ACM International Conference on Multimedia
Towards end-to-end explainable facial action unit recognition via vision- language joint learning. InProceedings of the 32nd ACM International Conference on Multimedia. 8189–8198
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.