REVIEW 4 major objections 5 minor 3 cited by
Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read MIntOOD combines weighted multimodal fusion, Dirichlet-synthesized pseudo-OOD examples, a cosine classifier, and contrastive learning so one model can classify known conversational intents and flag out-of-distribution utterances.
desk verdict Credible method, near-boundary-only OOD evaluation; needs a table fix and a cross-domain test before the headline claim is convincing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the pseudo-OOD feature, built as $z_{\mathrm{OOD}} = \sum_j \lambda_j z_{\mathrm{ID},j}$ with $\lambda$ sampled from a Dirichlet distribution with unit sum, drawn from $k=3$ ID features spanning at least two classes. It underpins all three objectives: a binary classifier gives coarse ID/OOD separation, a cosine classifier isolates directional information among intent classes, and a contrastive loss refines instance-level relations with dropout-changed duplicates as positive pairs. A weighted fusion network computes softmax-normalized modality weights per utterance and sums the encoded text, video, and audio representations; the Mahalanobis distance over per-class covariance then performs the final OOD scoring.
What would settle it
Build an OOD test set from a genuinely different source, such as utterances from a different conversational domain or language annotated by the same protocol, and compare MIntOOD's AUROC against baselines; if the margin disappears when the test OOD lies far from the convex combinations used in training, the pseudo-OOD proxy is the bottleneck. More directly, measure whether the model assigns high in-distribution scores to points on the inside of the ID convex hull; if it does, the binary head has learned to label hull-interior points as ID rather than semantic novelty.
Extended reading notes
Core claim
The central claim is that representations trained at several granularities on both in-distribution and synthesized out-of-distribution examples can simultaneously separate known intents and flag unseen ones. Pseudo-OOD features are made per modality as convex combinations of $k=3$ in-distribution features drawn from at least two classes, with weights sampled from a Dirichlet distribution; these are treated as negative examples in a binary ID/OOD head, while a cosine classifier separates the known classes and a contrastive loss organizes instances so that same-class samples cluster and pseudo-OOD samples stand apart. At inference, the Mahalanobis distance to per-class centroids in the fused feature space scores how out-of-distribution an utterance is. The paper reports that the full recipe beats every baseline on nearly all ID and OOD metrics across the three datasets, with the largest absolute gains on OOD detection.
Load-bearing premise
All of the OOD gains depend on the assumption that a convex mixture of a few in-distribution feature vectors, weighted by a Dirichlet draw, is close enough to a real out-of-scope utterance that training to reject it transfers to unseen inputs.
Editorial extensions
If this is right
- If the reported gains hold, a single multimodal model can serve both closed-world intent classification and open-world robustness, so dialogue systems need not degrade in accuracy to gain OOD awareness.
- Because pseudo-OOD data is generated in feature space from ID examples, the approach sidesteps the cost of collecting real out-of-scope utterances, which the paper calls prohibitively expensive.
- The ablation results indicate that the cosine classifier and Mahalanobis scoring are the main drivers of OOD performance, so future multimodal OOD detectors should not rely only on raw logits.
- The pseudo-OOD generation strategy is effective across two dialogue-act datasets and a fine-grained intent dataset, indicating the recipe transfers beyond 20-class intent taxonomies.
Reading between the lines
- Beyond the paper: a direct extension is to vary the number of mixed samples $k$ and the Dirichlet concentration $\alpha$ and measure AUROC on held-out OOD; the method's logic implies a sweet spot exists, with too-small mixtures indistinguishable from ID and too-large mixtures drifting beyond plausible OOD.
- Beyond the paper: the paper observes that gains are largest for the fine-grained 20-class MIntRec data, suggesting that the richness of the intent taxonomy may be what makes pseudo-OOD training pay off; this could be checked by subsampling MIntRec to fewer classes.
- Beyond the paper: because the pseudo-OOD operation happens in feature space before fusion, the same generation procedure could be dropped into other feature extractors or modality sets, though the paper does not demonstrate this transfer.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes MIntOOD, a method for joint multimodal intent classification and out-of-distribution (OOD) detection. The model extracts text, video, and audio features with pretrained backbones (BERT, Swin Transformer, WavLM), fuses them via a learned weighted fusion network, and generates pseudo-OOD samples as Dirichlet-weighted convex combinations of ID features from at least two classes (Eqs. 1–2). Training combines coarse-grained binary ID/OOD classification, a cosine classifier for ID classes, and contrastive learning for instance-level separation. At inference, OOD detection is performed with the Mahalanobis distance to the nearest class centroid (Eqs. 14–15). Experiments on MIntRec, MELD-DA, and IEMOCAP-DA report state-of-the-art ID classification and AUROC gains of roughly 3–10 absolute points over existing multimodal fusion baselines.
Significance. If the reported results hold, MIntOOD is a useful contribution to multimodal intent understanding: it provides a practical way to train OOD detectors without labeled OOD data, introduces a dynamic weighting fusion that benefits both tasks, and releases baselines and a new OOD benchmark for MIntRec. The ablation study (Table III) consistently shows that each component contributes on most datasets and metrics, and the public code and data availability are commendable. However, the significance is tempered by the evaluation protocol's limited out-of-domain scope and the absence of statistical uncertainty quantification, both of which directly affect the abstract's claim of generalization to 'unseen OOD data in real-world scenarios.'
major comments (4)
- [Sections IV-A and V-A (Tables I and II)] All three OOD test sets are drawn from the same distributional domain as the ID data: the MIntRec OOD utterances come from the same TV series (Superstore) as the ID data, and MELD-DA and IEMOCAP-DA treat the 'Others' dialogue-act label of the same corpus as OOD. Consequently, the reported AUROC improvements for the full MIntOOD (which trains only on pseudo-OOD generated by convex combinations of ID features) validate detection of near-manifold OOD, not the 'unseen OOD data in real-world scenarios' claimed in the abstract. A cross-domain experiment—for example, training the full MIntOOD on MIntRec ID data and testing on MELD-DA 'Others' or on an unrelated OOD collection—would directly substantiate the generalization claim. MIntOOD(R) does use real cross-domain OOD for training, but it does not test the pseudo-OOD generation itself in a cross-domain setting.
- [Table II, IEMOCAP-DA block] The MIntOOD(R) row for IEMOCAP-DA lists WF1=71.86, WP=72.59, F1=68.29, P=69.75, R=68.43, which are identical to the values in the MIntRec MIntOOD(R) row (only ACC differs). This is implausible for a different dataset and strongly suggests a copy-paste error. The authors should verify the reported averages and correct the table, since this error undermines confidence in the accuracy of the numerical results.
- [Section IV.D] All experimental results are reported as averages over five random seeds, but no standard deviations, confidence intervals, or significance tests are provided. Many of the claimed improvements are only 1–3 absolute points in AUROC or ACC (e.g., MIntRec AUROC 80.54 vs. 75.85 for the best baseline), and without variance estimates the reader cannot determine whether the differences are statistically reliable. The authors should report per-seed variability and, where possible, paired significance tests across the five seeds.
- [Section IV.A and Table I] The IEMOCAP-DA OOD test set contains only 70 samples. The near-perfect AUROC of 97.19 for MIntOOD on this set is therefore high-variance, and the large improvements in DER and FPR95 (44.85 and 46.57 points, respectively) should be interpreted with caution. A 95% confidence interval or a discussion of the small OOD sample size would make the strength of this result more assessable.
minor comments (5)
- [Section III-B, Eq. (2)] The notation C({yj}kj=1) is undefined; please define it as the set of unique class labels among the selected samples.
- [Figure 3] The figure appears to plot AUROC values, but the y-axis label is missing; please add a clear axis label and, if applicable, error bars.
- [Section V-A, last paragraph] The sentence 'with improvements of about 1% to 3%, 2% to 4%, and 2% to 40%, respectively' is ambiguous because the per-dataset mapping is not stated; specify which range corresponds to which dataset.
- [Section V-B, third paragraph] The sentence fragment 'the performance is comparable or even improves. .' contains a double period and is incomplete; rephrase to 'the performance is comparable or even slightly improves on IEMOCAP-DA.'
- [Related Work, Section II-A] Reference [43] (MIntRec2.0) is cited but not discussed in relation to the proposed OOD benchmark; clarify the relationship between MIntOOD and this closely related benchmark, including whether MIntRec2.0 contains an OOD split that could have been used for cross-domain evaluation.
Circularity Check
No circular derivation: pseudo-OOD is a training-time construct and the evaluation uses held-out real OOD with a Mahalanobis score, so the reported gains are not forced by construction.
full rationale
The claimed derivation chain is: generate pseudo-OOD features from convex combinations of ID features (Eq. 1-2), train coarse binary, cosine multi-class, and contrastive objectives on ID plus these synthetic samples, and then score held-out OOD with the Mahalanobis distance (Eq. 14-15). No step equates the test quantity with the training target. The pseudo-OOD data are used only for training; the OOD test sets (the 450 self-collected MIntRec utterances and the MELD-DA / IEMOCAP-DA 'Others' labels) are not used to construct Eq. 1-2, and the Mahalanobis statistics in Eq. 15 are computed from training ID features, not from pseudo-OOD or test OOD. MIntOOD (R) likewise trains on real OOD from a different dataset and is evaluated on held-out OOD, so no fitted parameter is renamed as a prediction. The paper does cite the authors' prior MIntRec [1] and TCL-MAP [11] work, but only as a dataset and a baseline, respectively; the central argument does not rest on an unverified self-citation or an imported uniqueness theorem. Two caveats belong in a correctness discussion rather than a circularity finding: the MIntRec OOD set comes from the same TV series as the ID data, so transfer to a genuinely different domain is not demonstrated, and the IEMOCAP-DA MIntOOD (R) ID metrics in Table II exactly duplicate the MIntRec row, which undermines confidence in that table but is not a circular reduction. MIntRec2.0 [43] is listed in the references but not discussed in the body, which is a completeness issue. Because the reported gains are empirical comparisons against external baselines on held-out OOD data, the paper's central claims are not equivalent to their inputs by construction.
Assumptions & free parameters
free parameters (5)
- k (number of mixed samples per pseudo-OOD example) =
3
- Dirichlet alpha =
(2, 0.7, 0.7)
- Cosine classifier scale gamma =
(16, 16, 32)
- Contrastive temperature tau =
(2, 1, 0.7)
- Learning rate =
(3e-5, 4e-6, 3e-6)
assumptions (4)
- domain assumption Convex combinations of ID features approximate real OOD features well enough for training to transfer.
- domain assumption The fused weighted-sum representation z_F preserves the information needed for both fine-grained ID classification and OOD detection.
- domain assumption Pretrained feature extractors (BERT, Swin Transformer, WavLM) provide suitable inputs for the task.
- domain assumption Per-class fused representations are approximately Gaussian with class-specific means and shared covariance for Mahalanobis scoring.
invented entities (1)
-
pseudo-OOD multimodal samples
independent evidence
Cite this review
Pith. "Pith review of Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding." pith.science (2026). https://pith.science/paper/ASBCE63G
@misc{pith2026241212453,
author = {Pith},
title = {Pith review of: Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding},
year = {2026},
howpublished = {\url{https://pith.science/paper/ASBCE63G}},
note = {Machine review of arXiv:2412.12453}
}
read the original abstract
Multimodal intent understanding is a significant research area that requires effective leveraging of multiple modalities to analyze human language. Existing methods face two main challenges in this domain. Firstly, they have limitations in capturing the nuanced and high-level semantics underlying complex in-distribution (ID) multimodal intents. Secondly, they exhibit poor generalization when confronted with unseen out-of-distribution (OOD) data in real-world scenarios. To address these issues, we propose a novel method for both ID classification and OOD detection (MIntOOD). We first introduce a weighted feature fusion network that models multimodal representations. This network dynamically learns the importance of each modality, adapting to multimodal contexts. To develop discriminative representations for both tasks, we synthesize pseudo-OOD data from convex combinations of ID data and engage in multimodal representation learning from both coarse-grained and fine-grained perspectives. The coarse-grained perspective focuses on distinguishing between ID and OOD binary classes, while the fine-grained perspective not only enhances the discrimination between different ID classes but also captures instance-level interactions between ID and OOD samples, promoting proximity among similar instances and separation from dissimilar ones. We establish baselines for three multimodal intent datasets and build an OOD benchmark. Extensive experiments on these datasets demonstrate that our method significantly improves OOD detection performance with a 3~10% increase in AUROC scores while achieving new state-of-the-art results in ID classification. Data and codes are available at https://github.com/thuiar/MIntOOD.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 3 Pith papers
-
Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
MMLA combines 61K multimodal utterances across six semantic dimensions; even the best fine-tuned multimodal LLM reaches only about 69% accuracy, exposing current limits in cognitive-level language understanding.
-
LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
LGSRR uses LLM-generated semantic descriptions and rankings to improve multimodal intent recognition, reporting SOTA results on MIntRec2.0 and IEMOCAP-DA with gains around 0.5-1.3%.
-
Deep Learning Approaches for Multimodal Intent Recognition: A Survey
A survey of deep learning methods for intent recognition, tracing the field from unimodal text, audio, vision, and EEG approaches to multimodal fusion, alignment, knowledge-augmented, and multi-task models.
Reference graph
Works this paper leans on
-
[1]
Mintrec: A new dataset for multimodal intent recognition,
H. Zhang, H. Xu, X. Wang, Q. Zhou, S. Zhao, and J. Teng, “Mintrec: A new dataset for multimodal intent recognition,” in Proc. 30th ACM Int. Conf. on Multimedia , 2022, pp. 1688–1697
work page 2022
-
[2]
Multimodal activity detection for natural interaction with virtual human,
K. Wang, S. Lian, H. Sang, W. Liu, Z. Liu, F. Shi, H. Deng, Z. Sun, and Z. Chen, “Multimodal activity detection for natural interaction with virtual human,” in Proc. 2023 IEEE Conf. Virtual Reality 3D User Interfaces Abstr. Workshops, 2023, pp. 671–672
work page 2023
-
[3]
Y . Fan, C. Wang, P. He, and Y . Hu, “Building multi-turn query interpreters for e-commercial chatbots with sparse-to-dense attentive modeling,” in Proc. 15th ACM Int. Conf. Web Search Data Mining , 2022, pp. 1577–1580
work page 2022
-
[4]
S. Paul, M. Sintek, V . K ¨epuska, M. Silaghi, and L. Robertson, “Intent based multimodal speech and gesture fusion for human-robot commu- nication in assembly situation,” in Proc. 2022 21st IEEE/CVF Int. Conf. Mach. Learn. Appl. , 2022, pp. 760–763
work page 2022
-
[5]
Towards a multimodal and context-aware framework for hu- man navigational intent inference,
Z. Zhang, “Towards a multimodal and context-aware framework for hu- man navigational intent inference,” in Proc. 2020 Int. Conf. Multimodal Interaction, 2020, pp. 738–742
work page 2020
-
[6]
Object affordance based multimodal fusion for natural human-robot interaction,
J. Mi, S. Tang, Z. Deng, M. G ¨orner, and J. Zhang, “Object affordance based multimodal fusion for natural human-robot interaction,” Cogn. Syst. Res., vol. 54, pp. 128–137, 2019
work page 2019
-
[7]
Multimodal transformer for unaligned multimodal language sequences,
Y .-H. H. Tsai, S. Bai, P. P. Liang, J. Z. Kolter, L.-P. Morency, and R. Salakhutdinov, “Multimodal transformer for unaligned multimodal language sequences,” in Proc. 57th Assoc. Comput. Linguist. , 2019, pp. 6558–6569
work page 2019
-
[8]
Misa: Modality-invariant and-specific representations for multimodal sentiment analysis,
D. Hazarika, R. Zimmermann, and S. Poria, “Misa: Modality-invariant and-specific representations for multimodal sentiment analysis,” in Proc. 28th ACM Int. Conf. on Multimedia , 2020, pp. 1122–1131
work page 2020
Show all 75 references
-
[9]
Integrating multimodal information in large pretrained transformers,
W. Rahman, M. K. Hasan, S. Lee, A. Zadeh, C. Mao, L.-P. Morency, and E. Hoque, “Integrating multimodal information in large pretrained transformers,” in Proc. 58th Assoc. Comput. Linguist. , 2020, pp. 2359– 2369
2020
-
[10]
Sdif-da: A shallow- to-deep interaction framework with data augmentation for multi-modal intent detection,
S. Huang, L. Qin, B. Wang, G. Tu, and R. Xu, “Sdif-da: A shallow- to-deep interaction framework with data augmentation for multi-modal intent detection,” in Proc. IEEE/CVF Int. Conf. Acoust., Speech Signal Process., 2024, pp. 10 206–10 210
2024
-
[11]
Token-level contrastive learning with modality-aware prompting for multimodal intent recognition,
Q. Zhou, H. Xu, H. Li, H. Zhang, X. Zhang, Y . Wang, and K. Gao, “Token-level contrastive learning with modality-aware prompting for multimodal intent recognition,” in Proc. AAAI Conf. Artif. Intell. , 2024, pp. 17 114–17 122
2024
-
[12]
Deep open intent classification with adaptive decision boundary,
H. Zhang, H. Xu, and T.-E. Lin, “Deep open intent classification with adaptive decision boundary,” in Proc. AAAI Conf. Artif. Intell. , 2021, pp. 14 374–14 382
2021
-
[13]
Learning discriminative representations and decision boundaries for open intent detection,
H. Zhang, H. Xu, S. Zhao, and Q. Zhou, “Learning discriminative representations and decision boundaries for open intent detection,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 31, pp. 1611– 1623, 2023
2023
-
[14]
An evaluation dataset for intent classification and out-of-scope prediction,
S. Larson, A. Mahendran, J. J. Peper, C. Clarke, A. Lee, P. Hill, J. K. Kummerfeld, K. Leach, M. A. Laurenzano, L. Tang, and J. Mars, “An evaluation dataset for intent classification and out-of-scope prediction,” in Proc. 2019 Conf. Empir. Methods Nat. Lang. Process. and 9th I...
2019
-
[15]
Generalized odin: Detecting out-of-distribution image without learning from out-of-distribution data,
Y . C. Hsu, Y . Shen, H. Jin, and Z. Kira, “Generalized odin: Detecting out-of-distribution image without learning from out-of-distribution data,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2020, pp. 10 951–10 960
2020
-
[16]
Energy-based out-of-distribution detection,
W. Liu, X. Wang, J. Owens, and Y . Li, “Energy-based out-of-distribution detection,” in Proc. Adv. Neural Inf. Process. Syst. , 2020, pp. 21 464– 21 475
2020
-
[17]
KNN-contrastive learning for out-of- domain intent classification,
Y . Zhou, P. Liu, and X. Qiu, “KNN-contrastive learning for out-of- domain intent classification,” in Proc. 60th Assoc. Comput. Linguist. , 2022, pp. 5129–5141
2022
-
[18]
Out-of-scope intent detection with self-supervision and discriminative training,
L.-M. Zhan, H. Liang, B. Liu, L. Fan, X.-M. Wu, and A. Y . Lam, “Out-of-scope intent detection with self-supervision and discriminative training,” in Proc. 59th Assoc. Comput. Linguist. , 2021, pp. 3521–3532
2021
-
[19]
Zero-shot out-of- distribution detection based on the pre-trained model clip,
S. Esmaeilpour, B. Liu, E. Robertson, and L. Shu, “Zero-shot out-of- distribution detection based on the pre-trained model clip,” inProc. AAAI Conf. Artif. Intell. , 2022, pp. 6568–6576
2022
-
[20]
Delving into out-of- distribution detection with vision-language representations,
Y . Ming, Z. Cai, J. Gu, Y . Sun, W. Li, and Y . Li, “Delving into out-of- distribution detection with vision-language representations,” Proc. Adv. Neural Inf. Process. Syst. , pp. 35 087–35 102, 2022
2022
-
[21]
A survey on spoken language understanding: Recent advances and new frontiers,
L. Qin, T. Xie, W. Che, and T. Liu, “A survey on spoken language understanding: Recent advances and new frontiers,” in Proc. Int. Joint Conf. Artif. Intell. , 2021, pp. 4577–4584
2021
-
[22]
Snips voice platform: an embedded spoken language understanding system for private-by-design voice interfaces,
A. Coucke, A. Saade, A. Ball, T. Bluche, A. Caulier, D. Leroy, C. Doumouro, T. Gisselbrecht, F. Caltagirone, T. Lavril et al. , “Snips voice platform: an embedded spoken language understanding system for private-by-design voice interfaces,” arXiv preprint arXiv:1805.10190 , 2018
2018 arXiv
-
[23]
Efficient intent detection with dual sentence encoders,
I. Casanueva, T. Tem ˇcinas, D. Gerz, M. Henderson, and I. Vuli ´c, “Efficient intent detection with dual sentence encoders,” in Proc. 2nd Workshop Nat. Lang. Process. for Conversational AI , 2020, pp. 38–45
2020
-
[24]
Bert for joint intent classification and slot filling,
Q. Chen, Z. Zhuo, and W. Wang, “Bert for joint intent classification and slot filling,” arXiv preprint arXiv:1902.10909 , 2019
1902 arXiv
-
[25]
Unified dialog model pre-training for task-oriented dialog understanding and generation,
W. He, Y . Dai, M. Yang, J. Sun, F. Huang, L. Si, and Y . Li, “Unified dialog model pre-training for task-oriented dialog understanding and generation,” in Proc. 45th Int. ACM SIGIR Conf. Res. Develop. Inf. Retrieval, 2022, pp. 187–200
2022
-
[26]
Discovering new intents via con- strained deep adaptive clustering with cluster refinement,
T.-E. Lin, H. Xu, and H. Zhang, “Discovering new intents via con- strained deep adaptive clustering with cluster refinement,” in Proc. AAAI Conf. Artif. Intell. , 2020, pp. 8360–8367
2020
-
[27]
Discovering new intents with deep aligned clustering,
H. Zhang, H. Xu, T.-E. Lin, and R. Lyu, “Discovering new intents with deep aligned clustering,” in Proc. AAAI Conf. Artif. Intell. , 2021, pp. 14 365–14 373
2021
-
[28]
A probabilistic framework for discov- ering new intents,
Y . Zhou, G. Quan, and X. Qiu, “A probabilistic framework for discov- ering new intents,” in Proc. 61st Assoc. Comput. Linguist. , 2023, pp. 3771–3784
2023
-
[29]
A clustering framework for unsupervised and semi-supervised new intent discovery,
H. Zhang, H. Xu, X. Wang, F. Long, and K. Gao, “A clustering framework for unsupervised and semi-supervised new intent discovery,” IEEE Trans. Knowl. Data Eng. , vol. 36, pp. 5468–5481, 2023
2023
-
[30]
Integrating text and image: Determining multimodal document intent 14 in Instagram posts,
J. Kruk, J. Lubin, K. Sikka, X. Lin, D. Jurafsky, and A. Divakaran, “Integrating text and image: Determining multimodal document intent 14 in Instagram posts,” in Proc. 2019 Conf. Empir. Methods Nat. Lang. Process. and 9th Int. Joint Conf. Nat. Lang. Process. , 2019, pp. 4622– 4632
2019
-
[31]
Multimodal marketing intent analysis for effective targeted advertising,
L. Zhang, J. Shen, J. Zhang, J. Xu, Z. Li, Y . Yao, and L. Yu, “Multimodal marketing intent analysis for effective targeted advertising,” IEEE Trans. Multimedia, vol. 24, pp. 1830–1843, 2021
2021
-
[32]
Emoint- trans: A multimodal transformer for identifying emotions and intents in social conversations,
G. V . Singh, M. Firdaus, A. Ekbal, and P. Bhattacharyya, “Emoint- trans: A multimodal transformer for identifying emotions and intents in social conversations,” IEEE/ACM Trans. Audio, Speech, Lang. Process., vol. 31, pp. 290–300, 2022
2022
-
[33]
Speech-text pre-training for spoken dialog understanding with explicit cross-modal alignment,
T. Yu, H. Gao, T.-E. Lin, M. Yang, Y . Wu, W. Ma, C. Wang, F. Huang, and Y . Li, “Speech-text pre-training for spoken dialog understanding with explicit cross-modal alignment,” in Proc. 61st Assoc. Comput. Linguist., 2023, pp. 7900–7913
2023
-
[34]
A deep multi-task model for dialogue act classification, intent detection and slot filling,
M. Firdaus, H. Golchha, A. Ekbal, and P. Bhattacharyya, “A deep multi-task model for dialogue act classification, intent detection and slot filling,” Cogn. Comput., vol. 13, pp. 626–645, 2021
2021
-
[35]
Switchboard: telephone speech corpus for research and development,
J. Godfrey, E. Holliman, and J. McDaniel, “Switchboard: telephone speech corpus for research and development,” in Proc. 1992 IEEE/CVF Int. Conf. Acoust., Speech, Signal Process. , 1992, pp. 517–520
1992
-
[36]
Dailydia- log: A manually labelled multi-turn dialogue dataset,
Y . Li, H. Su, X. Shen, W. Li, Z. Cao, and S. Niu, “Dailydia- log: A manually labelled multi-turn dialogue dataset,” arXiv preprint arXiv:1710.03957, 2017
2017 arXiv
-
[37]
Towards emotion- aided multi-modal dialogue act classification,
T. Saha, A. Patra, S. Saha, and P. Bhattacharyya, “Towards emotion- aided multi-modal dialogue act classification,” in Proc. 58th Assoc. Comput. Linguist., 2020, pp. 4361–4372
2020
-
[38]
Meld: A multimodal multi-party dataset for emotion recognition in conversations,
S. Poria, D. Hazarika, N. Majumder, G. Naik, E. Cambria, and R. Mihal- cea, “Meld: A multimodal multi-party dataset for emotion recognition in conversations,” in Proc. 57th Assoc. Comput. Linguist. , 2019, pp. 527–536
2019
-
[39]
Iemocap: Interactive emotional dyadic motion capture database,
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,” Lang. Resour. Eval., vol. 42, pp. 335– 359, 2008
2008
-
[40]
Ch-sims: A chinese multimodal sentiment analysis dataset with fine- grained annotation of modality,
W. Yu, H. Xu, F. Meng, Y . Zhu, Y . Ma, J. Wu, J. Zou, and K. Yang, “Ch-sims: A chinese multimodal sentiment analysis dataset with fine- grained annotation of modality,” in Proc. 58th Assoc. Comput. Linguist., 2020, pp. 3718–3727
2020
-
[41]
Mosi: multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos,
A. Zadeh, R. Zellers, E. Pincus, and L.-P. Morency, “Mosi: multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos,” arXiv preprint arXiv:1606.06259 , 2016
2016 arXiv
-
[42]
Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,
A. B. Zadeh, P. P. Liang, S. Poria, E. Cambria, and L.-P. Morency, “Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,” in Proc. 56th Assoc. Comput. Linguist., 2018, pp. 2236–2246
2018
-
[43]
MIntrec2.0: A large-scale benchmark dataset for multi- modal intent recognition and out-of-scope detection in conversations,
H. Zhang, X. Wang, H. Xu, Q. Zhou, J. Su, jinyue Zhao, W. Li, Y . Chen, and K. Gao, “MIntrec2.0: A large-scale benchmark dataset for multi- modal intent recognition and out-of-scope detection in conversations,” in Proc. 12th Int. Conf. Learn. Represent. , 2024
2024
-
[44]
Tensor fusion network for multimodal sentiment analysis,
A. Zadeh, M. Chen, S. Poria, E. Cambria, and L.-P. Morency, “Tensor fusion network for multimodal sentiment analysis,” in Proc. 2017 Conf. Empir. Methods Nat. Lang. Process. , 2017, pp. 1103–1114
2017
-
[45]
Efficient low-rank multimodal fusion with modality- specific factors,
Z. Liu, Y . Shen, V . B. Lakshminarasimhan, P. P. Liang, A. Bagher Zadeh, and L.-P. Morency, “Efficient low-rank multimodal fusion with modality- specific factors,” inProc. 56th Assoc. Comput. Linguist., 2018, pp. 2247– 2256
2018
-
[46]
Memory fusion network for multi-view sequential learning,
A. Zadeh, P. P. Liang, N. Mazumder, S. Poria, E. Cambria, and L.-P. Morency, “Memory fusion network for multi-view sequential learning,” in Proc. AAAI Conf. Artif. Intell. , 2018, pp. 5634–5641
2018
-
[47]
Improving multimodal fusion with hierarchical mutual information maximization for multimodal sentiment analysis,
W. Han, H. Chen, and S. Poria, “Improving multimodal fusion with hierarchical mutual information maximization for multimodal sentiment analysis,” in Proc. 2021 Conf. Empir. Methods Nat. Lang. Process., 2021, pp. 9180–9192
2021
-
[48]
Unimse: Towards unified multimodal sentiment analysis and emotion recognition,
G. Hu, T.-E. Lin, Y . Zhao, G. Lu, Y . Wu, and Y . Li, “Unimse: Towards unified multimodal sentiment analysis and emotion recognition,” inProc. 2022 Conf. Empir. Methods Nat. Lang. Process. , 2022, pp. 7837–7851
2022
-
[49]
Exploring the limits of transfer learning with a unified text-to-text transformer,
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res. , vol. 21, pp. 1–67, 2020
2020
-
[50]
Out-of-domain detection for natural language understanding in dialog systems,
Y . Zheng, G. Chen, and M. Huang, “Out-of-domain detection for natural language understanding in dialog systems,” IEEE/ACM Trans. Audio, Speech, Lang. Process., vol. 28, pp. 1198–1209, 2020
2020
-
[51]
Improving unsupervised out- of-domain detection through pseudo labeling and learning,
B. Lee, J. Kim, J. Park, and K.-A. Sohn, “Improving unsupervised out- of-domain detection through pseudo labeling and learning,” in Proc. Findings of the Assoc. Comput. Linguist.: EACL , 2023, pp. 1031–1041
2023
-
[52]
Enhancing the generalization for intent classification and out-of-domain detection in slu,
Y . Shen, Y .-C. Hsu, A. Ray, and H. Jin, “Enhancing the generalization for intent classification and out-of-domain detection in slu,” in Proc. 59th Assoc. Comput. Linguist. , 2021, pp. 2443–2453
2021
-
[53]
A closer look at few-shot out-of-distribution intent detection,
L.-M. Zhan, H. Liang, L. Fan, X.-M. Wu, and A. Y . Lam, “A closer look at few-shot out-of-distribution intent detection,” in Proc. 29th Int. Conf. Comput. Linguist. , 2022, pp. 451–460
2022
-
[54]
A simple unified framework for detecting out-of-distribution samples and adversarial attacks,
K. Lee, K. Lee, H. Lee, and J. Shin, “A simple unified framework for detecting out-of-distribution samples and adversarial attacks,” in Proc. Adv. Neural Inf. Process. Syst. , 2018, pp. 7166–7177
2018
-
[55]
Out-of-distribution detection with subspace techniques and probabilistic modeling of features,
I. Ndiour, N. Ahuja, and O. Tickoo, “Out-of-distribution detection with subspace techniques and probabilistic modeling of features,” arXiv preprint arXiv:2012.04250, 2020
2012 arXiv
-
[56]
Vim: Out-of-distribution with virtual-logit matching,
H. Wang, Z. Li, L. Feng, and W. Zhang, “Vim: Out-of-distribution with virtual-logit matching,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2022, pp. 4921–4930
2022
-
[57]
Swin transformer: Hierarchical vision transformer using shifted windows,
Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , 2021, pp. 9992– 10 002
2021
-
[58]
librosa: Audio and music signal analysis in python,
B. McFee, C. Raffel, D. Liang, D. P. Ellis, M. McVicar, E. Battenberg, and O. Nieto, “librosa: Audio and music signal analysis in python,” in Proc. 14th Python in Sci. Conf. , 2015, pp. 18–25
2015
-
[59]
Wavlm: Large-scale self-supervised pre- training for full stack speech processing,
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “Wavlm: Large-scale self-supervised pre- training for full stack speech processing,” IEEE J. Sel. Top. Signal Process., vol. 16, pp. 1505–1518, 2022
2022
-
[60]
Energy- based unknown intent detection with data manipulation,
Y . Ouyang, J. Ye, Y . Chen, X. Dai, S. Huang, and J. Chen, “Energy- based unknown intent detection with data manipulation,” in Findings of the Assoc. Comput. Linguist.: ACL-IJCNLP , 2021, pp. 2852–2861
2021
-
[61]
Learning to classify open intent via soft labeling and manifold mixup,
Z. Cheng, Z. Jiang, Y . Yin, C. Wang, and Q. Gu, “Learning to classify open intent via soft labeling and manifold mixup,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 30, pp. 635–645, 2022
2022
-
[62]
Bert: Pre-training of deep bidirectional transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proc. 2019 Conf. North Am. Chapter Assoc. Comput. Linguist.: Human Lang. Technol., 2019, pp. 4171–4186
2019
-
[63]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. Adv. Neural Inf. Process. Syst. , 2017, pp. 5998–6008
2017
-
[64]
Generalized out-of-distribution detection: A survey,
J. Yang, K. Zhou, Y . Li, and Z. Liu, “Generalized out-of-distribution detection: A survey,” Int. J. Comput. Vis. , vol. 132, pp. 5635–5662, 2024
2024
-
[65]
App: Adaptive prototypical pseudo-labeling for few-shot ood detection,
P. Wang, K. He, Y . Mou, X. Song, Y . Wu, J. Wang, Y . Xian, X. Cai, and W. Xu, “App: Adaptive prototypical pseudo-labeling for few-shot ood detection,” in Findings of the Assoc. Comput. Linguist.: EMNLP , 2023, pp. 3926–3939
2023
-
[66]
Low-shot learning with imprinted weights,
H. Qi, M. Brown, and D. G. Lowe, “Low-shot learning with imprinted weights,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2018, pp. 5822–5830
2018
-
[67]
Simcse: Simple contrastive learning of sentence embeddings,
T. Gao, X. Yao, and D. Chen, “Simcse: Simple contrastive learning of sentence embeddings,” in Proc. 2021 Conf. Empir. Methods Nat. Lang. Process., 2021, pp. 6894–6910
2021
-
[68]
Dialogue act modeling for automatic tagging and recognition of conversational speech,
A. Stolcke, K. Ries, N. Coccaro, E. Shriberg, R. Bates, D. Jurafsky, P. Taylor, R. Martin, C. V . Ess-Dykema, and M. Meteer, “Dialogue act modeling for automatic tagging and recognition of conversational speech,” Comput. Linguist., vol. 26, pp. 339–373, 2000
2000
-
[69]
A baseline for detecting misclassified and out-of-distribution examples in neural networks,
D. Hendrycks and K. Gimpel, “A baseline for detecting misclassified and out-of-distribution examples in neural networks,” in Proc. 5th Int. Conf. Learn. Represent. , 2017
2017
-
[70]
Enhancing the reliability of out-of- distribution image detection in neural networks,
S. Liang, Y . Li, and R. Srikant, “Enhancing the reliability of out-of- distribution image detection in neural networks,” in Proc. 6th Int. Conf. Learn. Represent., 2018
2018
-
[71]
Representation learning with contrastive predictive coding,
A. v. d. Oord, Y . Li, and O. Vinyals, “Representation learning with contrastive predictive coding,” arXiv preprint arXiv:1807.03748 , 2018
2018 arXiv
-
[72]
Out-of-domain detection based on generative adversarial network,
S. Ryu, S. Koo, H. Yu, and G. G. Lee, “Out-of-domain detection based on generative adversarial network,” in Proc. 2018 Conf. Empir. Methods Nat. Lang. Process. , 2018, pp. 714–718
2018
-
[73]
Transformers: State-of- the-art natural language processing,
T. Wolf, L. Debut, V . Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowiczet al., “Transformers: State-of- the-art natural language processing,” inProc. 2020 Conf. Empir. Methods Nat. Lang. Process.: Sys. Demonstrations , 2020, pp. 38–45
2020
-
[74]
Decoupled weight decay regularization,
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in Proc. 7th Int. Conf. Learn. Represent. , 2019
2019
-
[75]
Out-of-distribution detection using union of 1- dimensional subspaces,
A. Zaeemzadeh, N. Bisagno, Z. Sambugaro, N. Conci, N. Rah- navard, and M. Shah, “Out-of-distribution detection using union of 1- dimensional subspaces,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2021, pp. 9452–9461. 15 Hanlei Zhang received the B.S. degree from t...
2021
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.