{"id":"77780cda-3773-4568-8c7a-8ad27d2c146a","arxiv_id":"2508.09600","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A three-stage trained speech-to-speech chatbot with an explicit think step transfers paralinguistic understanding into empathetic responses, outperforming prior end-to-end spoken dialogue systems on a new LLM-scored benchmark.","lead":"OSUM-EChat is an open-source spoken chatbot trained to hear cues like age, gender, emotion, and coughs, then answer with empathy in speech. The paper also releases a 200,000-turn empathetic dialogue dataset and an evaluation benchmark, reporting that OSUM-EChat beats prior end-to-end spoken dialogue models on empathetic response scores.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"EChat-eval shares the DeepSeek/CosyVoice2 generation pipeline with EChat-200K training data; OSUM-EChat is evaluated in-distribution while baselines are out-of-distribution, so reported gains in Table 2 may reflect synthetic-distribution matching rather than empathetic ability.","rationale":"The paper's central claim is that OSUM-EChat outperforms end-to-end spoken dialogue models in empathy, but the only empathy benchmark that supports this is EChat-eval. The strongest independent evidence is human eval (Table 3), but it covers only OSUM-EChat, Qwen2.5-Omni, and Doubao, not the full baseline set, and it uses the same EChat-eval items—which are drawn from the same synthetic pipeline as training. The ablations (U-Driven / Dual Think) show relative gains within OSUM-EChat, supporting the training strategy's internal effectiveness, but they do not validate cross-model superiority. The Event column in Table 2 is never scored for baselines, and the keyword-matching rubric rewards the exact phrasing OSUM-EChat was trained to produce. The paper's own Limitation admits that the automated scoring differs substantially from human judgments. Together, these create a real risk that the headline margin is a benchmark artifact. This is not an accusation of fabrication; it is a missing control in the evaluation design. The proposed check—scoring the real-audio subset with blind human raters—would settle whether the advantage generalizes. If it does not, the paper's contribution would reduce to a data-generation/resource contribution, not a demonstrated empathetic ability advance. Credit is due for the speech understanding results on public test sets (Table 6, Table 8), which support the understanding-driven training claim in a distribution-independent way, and for the TTS/WER and VoiceBench results on standard benchmarks. So the paper has value independent of EChat-eval; the conditional is specifically on the empathy ranking. The reader's verdict of conditional acceptance is appropriate; my concern tightens the condition but does not change the verdict.","tokens_in":25550,"tokens_out":7502,"duration_ms":74902,"concrete_test":"Re-score the EChat-eval benchmark using only the approximately one-third real-audio entries (the portion not synthesized by the DeepSeek/CosyVoice2 pipeline), with human raters blind to model identity and reporting inter-annotator agreement. Compute the OSUM-EChat vs. Qwen2.5-Omni margin per dimension. If the margin collapses or reverses, the Table 2 advantage is an artifact of in-distribution training. If it persists with statistical significance, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"EChat-eval is constructed with 'the aforementioned data generation pipeline' (EChat-eval Bench), i.e., DeepSeek-generated queries/responses synthesized by CosyVoice2 (Appendix A.1), and OSUM-EChat's Stage-3 training uses the EChat-200K corpus drawn from the same pipeline (Table 1, 'Empathetic dialogue S2S'). Thus the model is evaluated in-distribution while all baselines in Table 2 (Qwen2.5-Omni, Freeze-Omni, GLM-4-Voice, etc.) are zero-shot on this distribution. The benchmark therefore measures how closely a model mimics the synthetic query→response mapping, not general empathetic ability. The Event metric is especially telling: scoring in Appendix A.2 is keyword matching (e.g., sigh:[sigh, ai, sighing...]; cough:[cough, coughing...]), and the EChat-200K construction prompts explicitly require the response to embed the caption label (e.g., 'Don't keep coughing—want some warm water?'). OSUM-EChat was trained to produce such references; baselines may express the same awareness with semantically equivalent but unlisted wording and be scored 0 or not scored. The paper's own Limitation concedes that emotion2vec label errors and ChatGPT-4o hallucinations make automated scores 'differ substantially from human evaluations.' Table 3's human eval covers only three models (OSUM-EChat, Qwen2.5-Omni, Doubao), with no inter-annotator agreement or confidence intervals, so it does not establish the full multi-model ranking claimed in the abstract. Without an out-of-distribution or human-scored test that includes the same baseline set, the headline superiority claim is confounded by training data overlap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes OSUM-EChat, an end-to-end empathetic spoken dialogue system built from a Whisper encoder, a Qwen2.5-3B LLM, and CosyVoice-based speech-token decoding. Training is organized into three stages—speech understanding (ASR+P), generation (TTS and S2S), and empathy—with a linguistic–paralinguistic dual think chain-of-thought inserted before response generation. The authors construct EChat-200K, a synthetic speech-to-speech empathetic dialogue corpus, and EChat-eval, a multidimensional benchmark scored by ChatGPT-4o with emotion2vec-extracted labels, plus a small human evaluation. The reported experiments (Tables 2–6) show consistent internal ablations and competitive linguistic/speech-understanding performance, and the headline table reports large advantages over open end-to-end baselines on EChat-eval.","tokens_in":25983,"tokens_out":7162,"duration_ms":78479,"significance":"If the empirical claims survive closer evaluation, the paper would make a useful contribution to resource-limited end-to-end empathetic spoken dialogue: the three-stage recipe is sensible, the dual-think mechanism is a concrete way to inject paralinguistic understanding into generation, and the open-source release (model, code, dataset, benchmark) is valuable for reproducibility. The ablations are consistent, and the public-test-set measurements (Table 6) are reassuring. However, the central comparative claim is not yet established because the main benchmark is generated by the same pipeline as the training set and scored by an LLM whose limitations the authors themselves acknowledge. The evidence is promising but requires an independent evaluation before the headline conclusion can be accepted.","major_comments":[{"comment":"The EChat-eval benchmark is 'constructed using the aforementioned data generation pipeline' (DeepSeek question/response generation plus CosyVoice2 synthesis), the same pipeline that produced the EChat-200K corpus used in Stage-3 empathy training (Table 1, 'Empathetic dialogue S2S'). OSUM-EChat is therefore evaluated in-distribution, while every baseline in Table 2 is zero-shot on this synthetic distribution. The reported gains (e.g., emotion 58.0 vs 44.8, multi-label 72.0 vs 63.7) may largely reflect matching the synthetic query→response distribution rather than general empathetic ability. The human validation in Table 3 covers only three models and does not include the main Table 2 baselines. Please add an out-of-distribution evaluation (e.g., human-scored natural dialogues, or public empathetic dialogue tests) covering all compared systems, or otherwise demonstrate that the benchmark i","section":"EChat-eval Bench / Appendix A.1; Table 2"},{"comment":"The sound-event dimension is scored by exact/surface keyword matching between the response and a fixed list (sigh: [sigh, ai, sighing, ...]; cough: [cough, coughing, ...]). The EChat-200K multi-label construction prompt explicitly instructs the generator to embed the caption label in the response (e.g., 'Don't keep coughing—want some warm water?'). Since OSUM-EChat is trained on responses produced under this instruction, it is specifically optimized to produce the lexical cues used by the scorer. A baseline that conveys the same awareness with different wording (e.g., 'Are you still under the weather?' after a cough) would receive a low/zero Event score. The 87.1 Event score is therefore not a credible quantification of empathetic event awareness. Use semantic or human scoring, and report the matching dictionary and failure cases.","section":"Appendix A.2, sound-event scoring; Table 2"},{"comment":"The paper concedes that emotion2vec label errors and ChatGPT-4o hallucinations cause automated scores to 'differ substantially from human evaluations,' yet the headline claim rests on the automated Table 2. The only human evaluation (Table 3) compares three systems, has dash entries for Event, and reports no number of annotators, per-item variance, confidence intervals, or inter-annotator agreement. Moreover, on the Emotion dimension the human score of OSUM-EChat (72.0) is lower than Doubao (78.0), so the 'outperforms' claim is already qualified in human evaluation. Statistical testing and a human study covering the main baselines and all dimensions are needed before the central claim can be accepted.","section":"Limitation; Tables 2-3"},{"comment":"The ablations U-Driven and Dual Think also rely on EChat-eval. Since the benchmark shares the training pipeline, the ablation gaps may be inflated by benchmark-specific behavior. The public-test-set results in Table 6 are more convincing; please complement the EChat-eval ablations with human ratings or an independent OOD test set to show that the two proposed components contribute to generalizable empathetic dialogue ability rather than only to synthetic-distribution matching.","section":"Main Results / Ablations; Table 2 lower half"}],"minor_comments":[{"comment":"OpenS2S appears in Table 2 but is not named in the contrastive models list. Also, the dash entries in the Event column should be explicitly defined (not evaluated vs. zero score) to avoid ambiguity.","section":"Contrastive Models; Table 2"},{"comment":"The emotion-label taxonomy is inconsistent: 'cheerful' appears in response labels in Figures 1 and 3, while the construction prompts use 'happy'; 'angry/fear/surprised' and 'anger/fear/surprise' also vary across prompts. This can affect emotion2vec mapping and the stability of automatic scoring.","section":"Appendix A.2 / Figures 1, 3"},{"comment":"The entry '0.2 Kh' is ambiguous: the text says approximately 200k conversations, but the column header defines Kh as thousands of hours. State total hours and number of conversations separately.","section":"Table 1 / EChat-200K"},{"comment":"The caption says 'WER%(↓) and CER%(↓) results' but the table only has test-zh and test-en columns; specify which metric is shown for each column and language.","section":"Table 5"},{"comment":"The reuse of LS2-S2S inside LS3 should be explicitly defined; as written it appears the Stage-2 loss term is added unweighted inside Stage 3, and the reader must infer the supervision format for the CoT tokens.","section":"Eq. (4)"}],"recommendation":"major_revision","confidential_remarks":"The paper's engineering contribution is real and the release is a service to the community. My main concern is benchmark circularity: EChat-eval is constructed from the same pipeline as EChat-200K, and the automatic scorer is an LLM with acknowledged hallucination issues. This is fixable in revision with a serious human/OOD evaluation, but as it stands the headline comparative claim is not yet supported. I would be willing to look at a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth a serious referee, but not because the headline comparison is convincing. The actual contribution is the understanding-driven training recipe and the EChat-200K corpus. Three-stage training that starts from an ASR+P understanding model and then adds generation and empathy is a sensible way to reduce data hunger for spoken dialogue, and the ablations in Table 2 are internally consistent: removing either the understanding-driven strategy or the dual-think mechanism drops scores across the board. The dataset, with single-label and multi-label paralinguistic annotations and a mix of synthetic and real audio, should be useful to the community. The authors also deserve credit for shipping code, data, and a stated open-source ecosystem, and they are honest in the Limitation section about the automatic scoring being flawed.\n\nThe soft spot is the evaluation, and it is a load-bearing one. EChat-eval is constructed with the same DeepSeek/CosyVoice2 pipeline as EChat-200K, and OSUM-EChat is trained on EChat-200K. So the headline results in Table 2 measure how well the model mimics the synthetic query-to-response mapping, while all baselines are zero-shot on that distribution. That is a genuine confound. The Event metric makes it worse: scoring is keyword matching, and the training prompts explicitly require the caption label to be embedded in the response (e.g., \"Don't keep coughing—want some warm water?\"). A baseline that conveys the same awareness with different wording gets scored zero, while OSUM-EChat is rewarded for doing exactly what it was trained to do. The stress-test note is correct on this. The human evaluation in Table 3 covers only three models, with no inter-annotator agreement or confidence intervals, so it does not rescue the full ranking. The paper's own claim that automated scores \"differ substantially from human evaluations\" undercuts the main table further.\n\nI would still send this to peer review. The training strategy and dataset are real contributions, the system is openly released, and the evaluation problem is fixable—rerun on a held-out set with out-of-distribution baselines, add human ratings for the full model set, and report variance. The authors seem aware of the limitations, which is a good sign. But the abstract's claim that OSUM-EChat \"outperforms other end-to-end spoken dialogue models\" is not supported by the current evidence. That needs to be tempered or re-proven.","headline":"Useful system and dataset, but the headline empathy claim is confounded by in-distribution evaluation; the right verdict is conditional acceptance after the eval is redone.","tokens_in":26543,"tokens_out":1114,"would_cite":true,"duration_ms":14898,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An end-to-end spoken dialogue system can become empathetic by reusing speech understanding: OSUM-EChat trains a 3B model in three stages and adds a linguistic–paralinguistic chain-of-thought, scoring 72.0 on multi-label EChat-eval and beati","keywords":["empathetic spoken dialogue","end-to-end speech language model","paralinguistic cues","chain-of-thought","speech-to-speech dialogue","EChat-200K","EChat-eval","resource-limited training"],"falsifier":"Replace ChatGPT-4o in EChat-eval with blind human raters scoring the same query speech, labels, and response audio, and also test on real unsynthesized emotional conversations. If the ranking between OSUM-EChat and the best compared end-to-end model flips on emotion or multi-label tasks—or if the human advantage disappears outside EChat-style synthesized data—then the 72.0 result is an artifact of the automatic scoring pipeline, not a stable empathetic ability.","tokens_in":25475,"feed_emoji":"🗣️","tokens_out":12586,"duration_ms":113401,"temperature":0.7,"pith_summary":"The paper aims to show that an end-to-end spoken dialogue system can produce empathetic replies without the huge proprietary dialogue datasets that current speech-language models depend on. The proposed system, OSUM-EChat, is trained in three stages—understanding, generation, empathy—so that paralinguistic perception learned by a speech understanding model is carried over into speech generation. Before each reply, the model runs a short chain-of-thought that names the user's words and their paralinguistic cues (age, gender, emotion, sound event), then speaks a response tailored to those cues. On the paper's new EChat-eval benchmark, OSUM-EChat scores 72.0 in multi-label empathetic scenarios, ahead of the end-to-end spoken models it is compared with, and ablations show both the training strategy and the chain-of-thought mechanism are needed for the gain. If these results hold, empathetic spoken assistants become feasible with limited data and open components.","feed_headline":"Spoken chatbot scores 72.0 on multi-label empathy benchmark","feed_subtitle":"Understanding-driven training plus a dual think step lets a 3B system read age, gender, emotion, and sound cues.","key_machinery":"Two coupled mechanisms carry the argument. The first is an understanding-driven spoken dialogue training strategy: three stages (understanding, generation, empathy) that first trains ASR plus paralinguistic multi-task recognition, then adds text-to-speech and speech-to-speech generation, then fine-tunes on the EChat-200K corpus. The second is a linguistic–paralinguistic dual think mechanism: a chain-of-thought enclosed in <think>...</think end> tokens in which the model restates the user's words and predicts age, gender, emotion, and sound-event labels before producing text and speech tokens. Together they let the model use speech-understanding knowledge as a scaffold for generation, so only","core_discovery":"On its own terms, the paper's central discovery is that paralinguistic understanding—recognizing who is speaking and how—can be deliberately transferred into spoken dialogue generation instead of being re-learned from millions of hours of speech conversation. OSUM-EChat starts with a pretrained speech understanding backbone, adds speech generation through a speech-token codec, and during an 'empathy' stage inserts a think block in which the model first restates the linguistic content and then predicts the paralinguistic labels before generating text and speech tokens. The trained 3B model reaches EChat-eval scores of 58.0 (emotion), 63.1 (age), 62.7 (gender), 87.1 (sound event), and 72.0 (mu","pith_inferences":["The paper leaves untested whether the dual think mechanism helps on natural human speech outside the EChat pipeline; a likely consequence of the design is that gains are largest when input speech is cleanly labeled, and the benchmark should be cross-checked on real call-center or podcast-style conversations.","Because paralinguistic labels are explicit tokens in the chain-of-thought, the architecture offers a natural control knob: a developer could override a predicted age, emotion, or event label at inference to force a different tone without retraining.","The same understanding-to-generation transfer may extend to other paralinguistic dimensions not in EChat-200K—such as dialect, health state, or social relationship—by adding label slots rather than redesigning the model."],"forward_implications":["If the paper is right, empathy in spoken assistants no longer requires closed, web-scale dialogue corpora; a 3B model plus the EChat-200K corpus reaches the top of the compared end-to-end systems on EChat-eval.","The think block provides a readable intermediate for paralinguistic reasoning, so a deployed assistant could in principle show or audit what cues it detected before responding.","The benchmark gives the field a common multi-dimensional yardstick (emotion, age, gender, sound events, multi-label) that can be reused to compare future end-to-end empathetic systems.","Ablation results imply that simply fine-tuning a speech-language model on empathetic data is not enough: both understanding-stage transfer and the dual think mechanism contribute to the outcome."],"supporting_citations":[{"why":"Supplies the OSUM speech understanding model and ASR+P multi-task strategy on which Stage 1 is built and which the dialogue model extends.","marker":"(Geng et al. 2025)"},{"why":"The pretrained Whisper encoder serves as OSUM-EChat's speech encoder for input audio features.","marker":"(Radford et al. 2023)"},{"why":"The Qwen2.5-3B-instruct model is the LLM decoder whose vocabulary is extended with speech tokens; it carries the linguistic knowledge.","marker":"(Yang et al. 2024)"},{"why":"CosyVoice provides the 4096 speech tokens and token2wav synthesis (flow matching plus HiFi-GAN) that let the LLM speak.","marker":"(Du et al. 2024a)"},{"why":"CosyVoice2 is used in the EChat-200K/EChat-eval pipeline to synthesize query and response speech from text and paralinguistic labels.","marker":"(Du et al. 2024b)"},{"why":"DeepSeek generates the empathetic query/response texts and paralinguistic labels that form the EChat-200K and EChat-eval content.","marker":"(Liu et al. 2024)"},{"why":"ChatGPT-4o is the core automated scorer in EChat-eval and also serves as a strong commercial baseline in the experiments.","marker":"(Hurst et al. 2024)"},{"why":"emotion2vec-Large extracts emotion labels from response speech, feeding the emotional-appropriateness part of EChat-eval scoring.","marker":"(Ma et al. 2024)"}],"fun_headline_variants":["Understanding-first training yields empathetic spoken responses","Dual think mechanism boosts empathy in spoken chatbots","OSUM-EChat: Empathy without massive dialogue datasets","Paralinguistic chain-of-thought improves spoken empathy","Chatbot decodes age, gender, emotion to speak empathetically"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that EChat-eval's automated scores—ChatGPT-4o judging emotion2vec-extracted labels—track what humans mean by empathy, and the paper's own limitation text concedes those automatic scores can diverge from human judgments; if they do not track true empathy, the headline advantage over other models may reflect fitting the synthetic benchmark rather than real empathetic ability.","fun_headline_variants_meta":{"raw":{"variants":["Understanding-first training yields empathetic spoken responses","Dual think mechanism boosts empathy in spoken chatbots","OSUM-EChat: Empathy without massive dialogue datasets","Paralinguistic chain-of-thought improves spoken empathy","Chatbot decodes age, gender, emotion to speak empathetically"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000363,"raw_usage":{"total_tokens":1833,"prompt_tokens":824,"completion_tokens":1009,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":568,"completion_tokens_details":{"reasoning_tokens":931}},"tokens_in":568,"tokens_out":1009,"duration_ms":11703,"temperature":1.0,"reasoning_tokens":931,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:56:37.786254+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replace ChatGPT-4o in EChat-eval with blind human raters scoring the same query speech, labels, and response audio, and also test on real unsynthesized emotional conversations. If the ranking between OSUM-EChat and the best compared end-to-end model flips on emotion or multi-label tasks—or if the human advantage disappears outside EChat-style synthesized data—then the 72.0 result is an artifact of the automatic scoring pipeline, not a stable empathetic ability.","supporting_citations":[],"review_version":1}