{"id":"05fb7fb8-a8d0-43eb-89fb-2e8b3a0cfc98","arxiv_id":"2507.09618","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"THAI-SER is the first sizeable Thai speech emotion recognition corpus, with 41.6 hours of acted and elicited speech, 27,854 utterances, and crowdsourced labels for five emotions.","lead":"This paper introduces THAI-SER, a new 41.6-hour speech emotion recognition dataset for Thai, with 27,854 utterances recorded by 200 actors in studio and Zoom settings. It provides crowdsourced emotion labels, reliability statistics, and baseline model results, filling a gap in non-Western SER resources.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The post-filter alpha of 0.692 is computed on the same data used to choose the 0.71 cutoff, so it is not a validated estimate of reliability for newly collected utterances.","rationale":"The corpus is a valuable and carefully documented resource, and the concern is not about fabrication or hidden data manipulation; the authors transparently report the raw alpha of 0.413 and the threshold-scanning procedure. The load-bearing issue is inferential: the central reliability claim uses the same data both to select the filtering cutoff and to compute the reported post-filter alpha. Since agreement filtering directly removes the items that depress alpha, the in-sample value of 0.692 is an optimistic estimate of what the recommended threshold would deliver on new data. The proposed five-fold held-out check is feasible because the corpus and annotation code are publicly released, and it would settle whether the 0.71 rule generalizes. If the held-out alpha stays above 0.667, the concern is resolved and the current claims stand. If it drops below, the paper should reframe the 0.692 figure as descriptive of the released subset rather than as a validated reliability guarantee for future filtered data. This does not change the reader's conditional verdict: the paper should be accepted only after the generalizability of the threshold is demonstrated or the claims are appropriately qualified.","tokens_in":26938,"tokens_out":4010,"duration_ms":47895,"concrete_test":"Run a five-fold threshold-validation experiment: split the 27,854 utterances into five folds stratified by recording session and emotion. For each fold, select the agreement threshold on the other four folds using the same scan as Section 4.2.1 (smallest threshold with alpha >= 0.667 under MASI distance), apply that threshold to the held-out fold, and compute Krippendorff's alpha and HRA on the retained held-out utterances. Report all five held-out values. If the minimum held-out alpha is below 0.667, the 0.71 threshold is overfit to the in-sample data and the headline reliability claim should be downgraded from prescriptive to descriptive.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.2.1 selects the agreement threshold 0.71 by scanning thresholds until Krippendorff's alpha crosses 0.667 on the full 27,854-utterance corpus (raw alpha 0.413). The headline statistics in Table 9 (alpha=0.692, HRA=0.772 on 14,182 retained utterances) are then computed on the same utterances that determined the cutoff. Because the filter preferentially removes low-consensus items, choosing the cutoff to hit the target alpha on the same data makes the reported reliability an in-sample, post hoc estimate. The paper's recommendation that users filter at 0.71 therefore carries an unvalidated guarantee: a user applying that threshold to new utterances, or even to a held-out portion of this corpus, may obtain alpha below 0.667. The claim that the corpus 'achieved an alpha score of 0.692, higher than a recommendation of 0.667' is thus not yet supported as a statement about the filtering procedure's generalizable reliability.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces THAI-SER, a new Thai speech emotion recognition corpus containing 27,854 utterances (41.61 hours) recorded from 200 professional actors in studio and Zoom environments, with both scripted and improvised sessions across five emotion categories. Annotations were collected via crowdsourcing with a multi-stage quality-control pipeline including pretests, gold utterances, consistency checks, and agreement-based filtering. The paper reports inter-annotator reliability (Krippendorff's alpha) and human recognition accuracy before and after filtering, analyzes reliability across demographics and recording conditions, and provides baseline and cross-corpus experiments. The corpus and experimental code are publicly released.","tokens_in":27154,"tokens_out":8398,"duration_ms":84294,"significance":"If the reliability claims hold, THAI-SER fills a clear gap as the first large-scale Thai (and a rare tonal-language) SER corpus, offering a valuable resource for speech emotion recognition in Southeast Asian languages. The paper's strengths include the size and diversity of the corpus, the inclusion of both controlled studio and realistic Zoom conditions, the unusually detailed documentation of the crowdsourcing quality-control pipeline, transparent reporting of raw unfiltered reliability metrics alongside filtered ones, and the provision of reproducible baseline code and cross-corpus benchmarks with public data release. The cross-corpus evaluation, including a size-matched pruning experiment with IEMOCAP, is a commendable effort at fair comparison.","major_comments":[{"comment":"The agreement threshold of 0.71 is selected by scanning thresholds on the full 27,854-utterance corpus until Krippendorff's alpha reaches the target of 0.667 (raw alpha 0.413). The headline reliability values (alpha = 0.692, HRA = 0.772 on the 14,182 retained utterances) are then computed on the same utterances that determined the cutoff. Because the filter preferentially removes low-consensus items, selecting the cutoff to achieve the target alpha on the same data makes the reported reliability an in-sample, post hoc estimate. The abstract's claim that the corpus 'achieved an alpha score of 0.692, higher than a recommendation of 0.667' is therefore not supported as a statement about the generalizable reliability of the filtering procedure: a user applying the 0.71 threshold to newly collected utterances, or to a held-out portion of this corpus, may obtain alpha below 0.667. Please validate the threshold on held-out data (e.g., select the threshold on a development split of sessions and report alpha on a test split), or provide bootstrap/confidence intervals for the filtered alpha, and temper the abstract and conclusion claims accordingly.","section":"Section 4.2.1, Figure 8, Table 9"}],"minor_comments":[{"comment":"The majority agreement formula is typeset ambiguously; the fraction appears to place the sum over annotators in the numerator and the sum over labels in the denominator, which is not mathematically well-defined. Please insert parentheses to make the intended expression clear, e.g., agreement(x_i) = max_k (1/N) * sum_{n=1}^N [ y_{nk} / (sum_{k'=1}^K y_{nk'}) ].","section":"Eq. (1)"},{"comment":"The HRA definition is not written as a statistic over utterances: the left-hand side is per-utterance while the sum runs over an index i up to N, without specifying whether N is the number of annotators or the number of utterances. Please rewrite as a corpus-level metric, e.g., HRA(D) = (1/|D|) sum_{x in D} 1(maj(y_x) = assigned(x)), and clarify the relationship to the values in Table 9.","section":"Eq. (6)"},{"comment":"THAI-SER is dated 2021 in Table 1, but the manuscript is from 2025; please correct the year or explain if 2021 refers to the collection period.","section":"Table 1"},{"comment":"The phrase 'optimal threshold' is used for 0.71; since the threshold is chosen to satisfy a target alpha on the current corpus, please qualify it as, e.g., 'the smallest threshold that achieves the target alpha on the present corpus,' to avoid implying a general optimum.","section":"Section 4.2.1"},{"comment":"The mapping of 'other' emotions to the five categories is subjective; in particular, mapping 'surprise' to 'happy' and 'calm' to 'neutral' may not be universally accepted. Please provide a brief justification or at least acknowledge the potential bias introduced by this mapping.","section":"Table 7"},{"comment":"The differences in alpha and HRA across gender, age, and recording conditions are reported without uncertainty estimates. Adding confidence intervals or significance tests would help readers gauge the strength of these descriptive comparisons.","section":"Section 4.2.2, Tables 9-10"}],"recommendation":"major_revision","confidential_remarks":"The main substantive risk is the circular threshold selection described in the major comment; this is fixable by adding a validation split or tempering the claims. Please also verify the novelty claim of being the first sizeable Thai SER corpus, since the authors only survey Thai ASR resources in Section 1.1 and a more thorough search for existing Thai affective speech databases would strengthen the contribution. The paper's broad scope (reliability analysis, baselines, cross-corpus) is appropriate for a dataset paper, but the reliability analysis needs the validation step before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"THAI-SER fills a real gap: it is the first sizeable Thai speech emotion corpus, with 41.6 hours, 200 actors, five emotions, scripted and improvised sessions, and studio plus Zoom conditions. The recording and annotation design is genuinely careful—pretests with trick questions, gold utterances, consistency checks, trustworthiness filtering, and manual handling of 'other' labels. The release is public, with code and baselines. Cross-corpus results against IEMOCAP, Emo-DB, and EMOVO are new and useful, even if rough. For the SER community and low-resource language work, this is a solid contribution.\n\nThe main soft spot is the reliability claim. The paper reports a post-filtering Krippendorff's alpha of 0.692, above the 0.667 recommendation, but Section 4.2.1 shows that the 0.71 agreement threshold was chosen by scanning the same 27,854 utterances until alpha crossed 0.667. The raw alpha is 0.413, and filtering keeps only about half the data (14,182 utterances). So the 0.692 is an in-sample, post hoc number, not a validated estimate of what a user would get applying that threshold to new data. The paper does show the threshold scan, which is honest, but the abstract's framing overstates the evidence. This is fixable: hold out a validation set, or present the alpha as a descriptive property of the filtered subset and soften the recommendation. As a resource paper, this should not block acceptance.\n\nTwo smaller issues: the actor demographics are internally inconsistent—Table 2 sums to 200 (112 female + 88 male) while the footnote says six actors are not described as male or female. And the cross-corpus tables lack error bars; the pruning procedure is a bit ad hoc. Both are minor.\n\nI would send this to review and ask the authors to address the threshold issue and the demographic inconsistency. The corpus itself is useful and the methods are transparent, so this deserves serious referee time.","headline":"THAI-SER is the first sizeable Thai SER corpus and worth having, but the headline alpha is partly a same-data threshold artifact; send to review with revisions.","tokens_in":27729,"tokens_out":3539,"would_cite":true,"duration_ms":38458,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces THAI-SER, the first sizeable Thai speech emotion recognition corpus, and shows that after filtering by a 0.71 agreement threshold its crowdsourced labels reach a Krippendorff's alpha of 0.692 and a human recognition…","keywords":["speech emotion recognition","Thai language","acted speech corpus","elicited speech","crowdsourced annotation","inter-annotator reliability","Krippendorff's alpha","tonal language"],"falsifier":"Take a fresh held-out batch of THAI-SER utterances, or re-annotate a random subset with the same pretest, gold, and consistency checks, then apply the fixed 0.71 filter and report alpha and human recognition accuracy on that batch alone; if alpha falls below 0.667, or accuracy falls well below 0.772, the reliability numbers are not portable to new data.","tokens_in":26766,"feed_emoji":"🎭","tokens_out":11849,"duration_ms":108626,"temperature":0.7,"pith_summary":"THAI-SER is the first sizeable Thai speech emotion recognition corpus, built from 41 hours 36 minutes of audio divided into 27,854 utterances by 200 professional actors, covering neutral, angry, happy, sad, and frustrated speech. The recordings mix scripted and improvised sessions and studio and Zoom environments, so the corpus can support both controlled acted-emotion research and out-of-domain testing for online meetings. The paper's central claim is that a crowdsourced annotation pipeline, with a pretest, gold and consistency questions, and an agreement-based filter at 0.71, yields labels reliable enough for research: the filtered corpus reaches a Krippendorff's alpha of 0.692, above the 0.667 recommendation, and a human recognition accuracy of 0.772. The paper also provides speaker-independent k-fold baselines and cross-corpus comparisons with IEMOCAP, Emo-DB, and EMOVO. A reader should care because Thai is a tonal language whose prosodic emotion cues may differ from the Western languages that dominate existing SER corpora.","feed_headline":"First sizeable Thai emotion-speech corpus passes reliability bar","feed_subtitle":"It totals 41.6 hours and 27,854 utterances from 200 actors, with 77.2 percent human recognition accuracy.","key_machinery":"The load-bearing machinery is the annotation-quality pipeline rather than a single algorithm. Each utterance receives 3 to 8 crowdsourced labels; annotator trustworthiness is gated by gold utterances and duplicated consistency utterances in addition to a pretest with a trick question; per-utterance majority agreement is computed as the proportion of annotators selecting the majority emotion; then a cutoff of 0.71 is chosen by scanning agreement thresholds until the corpus-level Krippendorff's alpha, computed with the MASI distance for set-valued labels, reaches at least 0.667. The corpus design carries part of the argument as well: fixed emotion-neutral sentences for scripted takes strip out lexical context, improvised dyadic scenarios elicit more natural speech, and the two recording environments, studio and Zoom, create an in-domain and out-of-domain split used throughout the experiments.","core_discovery":"On its own terms, the paper establishes that a large Thai acted-and-elicited corpus can be annotated to research-grade reliability through a carefully gated crowdsourcing pipeline: the raw corpus has low agreement (alpha 0.413), but removing the 13,672 utterances with majority agreement below 0.71 leaves 14,182 utterances with alpha 0.692 and human recognition accuracy 0.772. The paper further shows that reliability is not uniform across the corpus: scripted high-intensity performances are the easiest for humans to recognize (0.883 human recognition accuracy), low-intensity scripted performances are the hardest (0.690), improvised sessions reach higher inter-annotator reliability than scripted sessions overall, and frustrated is the emotion most often confused with angry and sad. The authors present these differences as evidence that emotional intensity and acting style are design variables rather than noise, and that the filtered THAI-SER can serve as a benchmark and a cross-corpus testbed for speech emotion recognition.","pith_inferences":["If the 0.71 threshold generalizes beyond this corpus, the same pretest-plus-gold-plus-consistency annotation pipeline becomes a portable recipe for other under-resourced tonal languages, and its transferability could be tested by applying it to a small pilot corpus in a related language without re-tuning the cutoff.","Because THAI-SER carries up to eight annotations per utterance, it offers soft-label distributions richer than IEMOCAP's fixed three-annotator design; a test the paper does not run is whether training on those soft labels, or on ambiguous samples in curriculum order, closes part of the gap with human recognition accuracy.","The confusion between frustrated, angry, and sad suggests a hierarchical label scheme, for example coarse arousal and valence before fine emotion, might outperform flat five-way classification; this is directly testable on the released corpus.","The Zoom sessions could also be used to separate the causes of the out-of-domain performance drop by taking clean studio audio, passing it through the same Zoom encoding path, and comparing model accuracy before and after that controlled degradation."],"forward_implications":["Filtered THAI-SER gives Thai speech emotion recognition its first speaker-independent benchmark: with the four basic emotions, the provided CNN+LSTM baseline reaches 67.34% weighted and 62.61% unweighted accuracy, and adding frustration lowers both to 59.80% and 57.81%.","Scripted-only training beats improvised-only and mixed training on THAI-SER (73.99% weighted accuracy for four emotions), the opposite of the usual IEMOCAP result, so acting style should be reported and tuned per corpus.","Studio-trained models drop sharply on Zoom recordings, to roughly 46 to 56 percent weighted accuracy, making the Zoom split a ready-made robustness benchmark for future speech emotion recognition work.","High-intensity scripted emotions are much easier to recognize than low-intensity ones (0.883 versus 0.690 human recognition accuracy), so corpus builders can treat intensity as a design lever rather than an incidental property of acted speech.","Cross-corpus results position THAI-SER as a useful source language: THAI-SER-trained models beat IEMOCAP-trained models on Emo-DB and EMOVO even after both corpora are pruned to matched hours and speakers."],"supporting_citations":[{"why":"Establishes Krippendorff's alpha as the inter-annotator reliability metric and the 0.667 threshold the corpus is judged against.","marker":"Krippendorff (2004)"},{"why":"Defines the MASI distance used to compute alpha on set-valued emotion labels.","marker":"Passonneau (2006)"},{"why":"Provides the acted-plus-elicited dyadic design for the corpus and the IEMOCAP dataset used in cross-corpus tests.","marker":"Busso et al. (2008)"},{"why":"Demonstrates the large-scale crowdsourced annotation model and supplies CREMA-D as a comparison corpus.","marker":"Cao et al. (2014)"},{"why":"Supplies the speaker-independent k-fold cross-validation protocol and the CNN+LSTM architecture used for all baselines.","marker":"Etienne et al. (2018)"},{"why":"Reports the IEMOCAP finding that improvised-only training works best, which this paper tests and contrasts with its scripted-only result.","marker":"Neumann and Vu (2017)"},{"why":"Supplies Emo-DB, a German acted corpus used as a cross-corpus evaluation target.","marker":"Burkhardt et al. (2005)"},{"why":"Shows that elicited emotional speech can be recorded over video chat, the precedent for the Zoom sessions and out-of-domain evaluation.","marker":"Kossaifi et al. (2021)"}],"fun_headline_variants":["Thai SER: 27,854 utterances, 41.6 hours, alpha 0.692","Thai emotion corpus: filtered to 14,182 utterances for research","Thai acting corpus reaches alpha 0.692 after crowd filtering","Public Thai SER corpus offers 41.6 hours of emotion data","Thai SER: 200 actors, 5 emotions, crowd-verified alpha 0.692"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 0.71 cutoff was picked by testing different cutoffs on the same corpus until a standard reliability score rose above 0.667, and the headline reliability figures depend on that cutoff continuing to work when new utterances are annotated.","fun_headline_variants_meta":{"raw":{"variants":["Thai SER: 27,854 utterances, 41.6 hours, alpha 0.692","Thai emotion corpus: filtered to 14,182 utterances for research","Thai acting corpus reaches alpha 0.692 after crowd filtering","Public Thai SER corpus offers 41.6 hours of emotion data","Thai SER: 200 actors, 5 emotions, crowd-verified alpha 0.692"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000945,"raw_usage":{"total_tokens":4065,"prompt_tokens":1004,"completion_tokens":3061,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":620,"completion_tokens_details":{"reasoning_tokens":2958}},"tokens_in":620,"tokens_out":3061,"duration_ms":24393,"temperature":1.0,"reasoning_tokens":2958,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:51:53.004762+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a fresh held-out batch of THAI-SER utterances, or re-annotate a random subset with the same pretest, gold, and consistency checks, then apply the fixed 0.71 filter and report alpha and human recognition accuracy on that batch alone; if alpha falls below 0.667, or accuracy falls well below 0.772, the reliability numbers are not portable to new data.","supporting_citations":[{"cited_title":"APACrefauthors \\ 2004","cited_arxiv_id":null,"evidence_quote":"Establishes Krippendorff's alpha as the inter-annotator reliability metric and the 0.667 threshold the corpus is judged against."},{"cited_title":"APACrefauthors \\ 2006","cited_arxiv_id":null,"evidence_quote":"Defines the MASI distance used to compute alpha on set-valued emotion labels."},{"cited_title":", Cooper , D.G","cited_arxiv_id":null,"evidence_quote":"Demonstrates the large-scale crowdsourced annotation model and supplies CREMA-D as a comparison corpus."},{"cited_title":"\\ Vu, N.T","cited_arxiv_id":null,"evidence_quote":"Reports the IEMOCAP finding that improvised-only training works best, which this paper tests and contrasts with its scripted-only result."},{"cited_title":", Paeschke, A","cited_arxiv_id":null,"evidence_quote":"Supplies Emo-DB, a German acted corpus used as a cross-corpus evaluation target."},{"cited_title":", Walecki , R","cited_arxiv_id":null,"evidence_quote":"Shows that elicited emotional speech can be recorded over video chat, the precedent for the Zoom sessions and out-of-domain evaluation."}],"review_version":1}