Pith. sign in

REVIEW 4 major objections 5 minor 3 cited by

Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read MIntOOD combines weighted multimodal fusion, Dirichlet-synthesized pseudo-OOD examples, a cosine classifier, and contrastive learning so one model can classify known conversational intents and flag out-of-distribution utterances.

desk verdict Credible method, near-boundary-only OOD evaluation; needs a table fix and a cross-domain test before the headline claim is convincing. read the letter →

arxiv 2412.12453 v2 pith:ASBCE63G submitted 2024-12-17 cs.MM

classification cs.MM
keywords multimodalintentunderstandingout-of-distributiondetectionpseudo-OODgenerationweightedfeaturefusioncosineclassifiercontrastivelearningDirichletmixingdialogueactclassification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Multimodal intent understanding asks a system to read what a speaker means from text, video, and audio, and to know when an utterance belongs to none of the trained intents. This paper proposes MIntOOD as one method for both jobs: it fuses the three modalities with per-utterance learned weights, and trains on original in-distribution data plus pseudo-OOD features synthesized by mixing in-distribution features with Dirichlet-sampled coefficients. The paper's reported results are state-of-the-art in-distribution accuracy on MIntRec, MELD-DA, and IEMOCAP-DA, and out-of-distribution AUROC gains of about 3–10 points over the best baselines. If correct, the contribution is evidence that open-world robustness and closed-world accuracy are not in conflict for multimodal intent systems.

What carries the argument

The load-bearing object is the pseudo-OOD feature, built as $z_{\mathrm{OOD}} = \sum_j \lambda_j z_{\mathrm{ID},j}$ with $\lambda$ sampled from a Dirichlet distribution with unit sum, drawn from $k=3$ ID features spanning at least two classes. It underpins all three objectives: a binary classifier gives coarse ID/OOD separation, a cosine classifier isolates directional information among intent classes, and a contrastive loss refines instance-level relations with dropout-changed duplicates as positive pairs. A weighted fusion network computes softmax-normalized modality weights per utterance and sums the encoded text, video, and audio representations; the Mahalanobis distance over per-class covariance then performs the final OOD scoring.

What would settle it

Build an OOD test set from a genuinely different source, such as utterances from a different conversational domain or language annotated by the same protocol, and compare MIntOOD's AUROC against baselines; if the margin disappears when the test OOD lies far from the convex combinations used in training, the pseudo-OOD proxy is the bottleneck. More directly, measure whether the model assigns high in-distribution scores to points on the inside of the ID convex hull; if it does, the binary head has learned to label hull-interior points as ID rather than semantic novelty.

Watch

Extended reading notes

Core claim

The central claim is that representations trained at several granularities on both in-distribution and synthesized out-of-distribution examples can simultaneously separate known intents and flag unseen ones. Pseudo-OOD features are made per modality as convex combinations of $k=3$ in-distribution features drawn from at least two classes, with weights sampled from a Dirichlet distribution; these are treated as negative examples in a binary ID/OOD head, while a cosine classifier separates the known classes and a contrastive loss organizes instances so that same-class samples cluster and pseudo-OOD samples stand apart. At inference, the Mahalanobis distance to per-class centroids in the fused feature space scores how out-of-distribution an utterance is. The paper reports that the full recipe beats every baseline on nearly all ID and OOD metrics across the three datasets, with the largest absolute gains on OOD detection.

Load-bearing premise

All of the OOD gains depend on the assumption that a convex mixture of a few in-distribution feature vectors, weighted by a Dirichlet draw, is close enough to a real out-of-scope utterance that training to reject it transfers to unseen inputs.

Editorial extensions

If this is right

  • If the reported gains hold, a single multimodal model can serve both closed-world intent classification and open-world robustness, so dialogue systems need not degrade in accuracy to gain OOD awareness.
  • Because pseudo-OOD data is generated in feature space from ID examples, the approach sidesteps the cost of collecting real out-of-scope utterances, which the paper calls prohibitively expensive.
  • The ablation results indicate that the cosine classifier and Mahalanobis scoring are the main drivers of OOD performance, so future multimodal OOD detectors should not rely only on raw logits.
  • The pseudo-OOD generation strategy is effective across two dialogue-act datasets and a fine-grained intent dataset, indicating the recipe transfers beyond 20-class intent taxonomies.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: a direct extension is to vary the number of mixed samples $k$ and the Dirichlet concentration $\alpha$ and measure AUROC on held-out OOD; the method's logic implies a sweet spot exists, with too-small mixtures indistinguishable from ID and too-large mixtures drifting beyond plausible OOD.
  • Beyond the paper: the paper observes that gains are largest for the fine-grained 20-class MIntRec data, suggesting that the richness of the intent taxonomy may be what makes pseudo-OOD training pay off; this could be checked by subsampling MIntRec to fewer classes.
  • Beyond the paper: because the pseudo-OOD operation happens in feature space before fusion, the same generation procedure could be dropped into other feature extractors or modality sets, though the paper does not demonstrate this transfer.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper proposes MIntOOD, a method for joint multimodal intent classification and out-of-distribution (OOD) detection. The model extracts text, video, and audio features with pretrained backbones (BERT, Swin Transformer, WavLM), fuses them via a learned weighted fusion network, and generates pseudo-OOD samples as Dirichlet-weighted convex combinations of ID features from at least two classes (Eqs. 1–2). Training combines coarse-grained binary ID/OOD classification, a cosine classifier for ID classes, and contrastive learning for instance-level separation. At inference, OOD detection is performed with the Mahalanobis distance to the nearest class centroid (Eqs. 14–15). Experiments on MIntRec, MELD-DA, and IEMOCAP-DA report state-of-the-art ID classification and AUROC gains of roughly 3–10 absolute points over existing multimodal fusion baselines.

Significance. If the reported results hold, MIntOOD is a useful contribution to multimodal intent understanding: it provides a practical way to train OOD detectors without labeled OOD data, introduces a dynamic weighting fusion that benefits both tasks, and releases baselines and a new OOD benchmark for MIntRec. The ablation study (Table III) consistently shows that each component contributes on most datasets and metrics, and the public code and data availability are commendable. However, the significance is tempered by the evaluation protocol's limited out-of-domain scope and the absence of statistical uncertainty quantification, both of which directly affect the abstract's claim of generalization to 'unseen OOD data in real-world scenarios.'

major comments (4)
  1. [Sections IV-A and V-A (Tables I and II)] All three OOD test sets are drawn from the same distributional domain as the ID data: the MIntRec OOD utterances come from the same TV series (Superstore) as the ID data, and MELD-DA and IEMOCAP-DA treat the 'Others' dialogue-act label of the same corpus as OOD. Consequently, the reported AUROC improvements for the full MIntOOD (which trains only on pseudo-OOD generated by convex combinations of ID features) validate detection of near-manifold OOD, not the 'unseen OOD data in real-world scenarios' claimed in the abstract. A cross-domain experiment—for example, training the full MIntOOD on MIntRec ID data and testing on MELD-DA 'Others' or on an unrelated OOD collection—would directly substantiate the generalization claim. MIntOOD(R) does use real cross-domain OOD for training, but it does not test the pseudo-OOD generation itself in a cross-domain setting.
  2. [Table II, IEMOCAP-DA block] The MIntOOD(R) row for IEMOCAP-DA lists WF1=71.86, WP=72.59, F1=68.29, P=69.75, R=68.43, which are identical to the values in the MIntRec MIntOOD(R) row (only ACC differs). This is implausible for a different dataset and strongly suggests a copy-paste error. The authors should verify the reported averages and correct the table, since this error undermines confidence in the accuracy of the numerical results.
  3. [Section IV.D] All experimental results are reported as averages over five random seeds, but no standard deviations, confidence intervals, or significance tests are provided. Many of the claimed improvements are only 1–3 absolute points in AUROC or ACC (e.g., MIntRec AUROC 80.54 vs. 75.85 for the best baseline), and without variance estimates the reader cannot determine whether the differences are statistically reliable. The authors should report per-seed variability and, where possible, paired significance tests across the five seeds.
  4. [Section IV.A and Table I] The IEMOCAP-DA OOD test set contains only 70 samples. The near-perfect AUROC of 97.19 for MIntOOD on this set is therefore high-variance, and the large improvements in DER and FPR95 (44.85 and 46.57 points, respectively) should be interpreted with caution. A 95% confidence interval or a discussion of the small OOD sample size would make the strength of this result more assessable.
minor comments (5)
  1. [Section III-B, Eq. (2)] The notation C({yj}kj=1) is undefined; please define it as the set of unique class labels among the selected samples.
  2. [Figure 3] The figure appears to plot AUROC values, but the y-axis label is missing; please add a clear axis label and, if applicable, error bars.
  3. [Section V-A, last paragraph] The sentence 'with improvements of about 1% to 3%, 2% to 4%, and 2% to 40%, respectively' is ambiguous because the per-dataset mapping is not stated; specify which range corresponds to which dataset.
  4. [Section V-B, third paragraph] The sentence fragment 'the performance is comparable or even improves. .' contains a double period and is incomplete; rephrase to 'the performance is comparable or even slightly improves on IEMOCAP-DA.'
  5. [Related Work, Section II-A] Reference [43] (MIntRec2.0) is cited but not discussed in relation to the proposed OOD benchmark; clarify the relationship between MIntOOD and this closely related benchmark, including whether MIntRec2.0 contains an OOD split that could have been used for cross-domain evaluation.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: pseudo-OOD is a training-time construct and the evaluation uses held-out real OOD with a Mahalanobis score, so the reported gains are not forced by construction.

full rationale

The claimed derivation chain is: generate pseudo-OOD features from convex combinations of ID features (Eq. 1-2), train coarse binary, cosine multi-class, and contrastive objectives on ID plus these synthetic samples, and then score held-out OOD with the Mahalanobis distance (Eq. 14-15). No step equates the test quantity with the training target. The pseudo-OOD data are used only for training; the OOD test sets (the 450 self-collected MIntRec utterances and the MELD-DA / IEMOCAP-DA 'Others' labels) are not used to construct Eq. 1-2, and the Mahalanobis statistics in Eq. 15 are computed from training ID features, not from pseudo-OOD or test OOD. MIntOOD (R) likewise trains on real OOD from a different dataset and is evaluated on held-out OOD, so no fitted parameter is renamed as a prediction. The paper does cite the authors' prior MIntRec [1] and TCL-MAP [11] work, but only as a dataset and a baseline, respectively; the central argument does not rest on an unverified self-citation or an imported uniqueness theorem. Two caveats belong in a correctness discussion rather than a circularity finding: the MIntRec OOD set comes from the same TV series as the ID data, so transfer to a genuinely different domain is not demonstrated, and the IEMOCAP-DA MIntOOD (R) ID metrics in Table II exactly duplicate the MIntRec row, which undermines confidence in that table but is not a circular reduction. MIntRec2.0 [43] is listed in the references but not discussed in the body, which is a completeness issue. Because the reported gains are empirical comparisons against external baselines on held-out OOD data, the paper's central claims are not equivalent to their inputs by construction.

Assumptions & free parameters 5 free parameters · 4 assumptions · 1 invented entities

The central training signal is the assumption that convex combinations of existing ID features represent unseen OOD data. This is not proven, only indirectly tested through real OOD benchmarks. Hyperparameters are tuned per dataset, so reported numbers include fitting choices, but no parameter is fitted to the OOD test labels.

free parameters (5)
  • k (number of mixed samples per pseudo-OOD example) = 3
    Number of ID embeddings mixed to form each pseudo-OOD vector. Hand-selected; no sensitivity analysis is reported.
  • Dirichlet alpha = (2, 0.7, 0.7)
    Controls pseudo-OOD diversity; tuned per dataset and reported in Section IV-D.
  • Cosine classifier scale gamma = (16, 16, 32)
    Output scaling for the cosine classifier; tuned per dataset and reported in Section IV-D.
  • Contrastive temperature tau = (2, 1, 0.7)
    Temperature in contrastive losses; tuned per dataset and reported in Section IV-D.
  • Learning rate = (3e-5, 4e-6, 3e-6)
    AdamW learning rate per dataset; tuned on validation ID data and reported in Section IV-D.
assumptions (4)
  • domain assumption Convex combinations of ID features approximate real OOD features well enough for training to transfer.
    Section III-B, Eq. (1)-(2). The entire pseudo-OOD training signal depends on this proxy being good enough to transfer to real unseen utterances; the paper only tests this indirectly.
  • domain assumption The fused weighted-sum representation z_F preserves the information needed for both fine-grained ID classification and OOD detection.
    Section III-C, Eq. (8). Video and audio are mean-pooled and linearly projected to text space, discarding temporal structure before fusion.
  • domain assumption Pretrained feature extractors (BERT, Swin Transformer, WavLM) provide suitable inputs for the task.
    Section III-B. The method inherits any biases or errors in these backbones; no joint end-to-end fine-tuning of the extractors is described beyond the encoders.
  • domain assumption Per-class fused representations are approximately Gaussian with class-specific means and shared covariance for Mahalanobis scoring.
    Section III-E, Eqs. (14)-(15). All methods are scored the same way, but the Gaussian assumption is not validated.
invented entities (1)
  • pseudo-OOD multimodal samples independent evidence
    purpose: Synthetic training examples that stand in for unknown, out-of-distribution utterances so the model can learn ID/OOD boundaries without labeled OOD data.
    The paper tests the construct on real OOD samples not used to generate it (annotated MIntRec OOD and Others-class dialogue acts), so its utility is falsifiable, though the realism of the proxy itself is not directly measured.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding." pith.science (2026). https://pith.science/paper/ASBCE63G

@misc{pith2026241212453,
  author       = {Pith},
  title        = {Pith review of: Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ASBCE63G}},
  note         = {Machine review of arXiv:2412.12453}
}
read the original abstract

Multimodal intent understanding is a significant research area that requires effective leveraging of multiple modalities to analyze human language. Existing methods face two main challenges in this domain. Firstly, they have limitations in capturing the nuanced and high-level semantics underlying complex in-distribution (ID) multimodal intents. Secondly, they exhibit poor generalization when confronted with unseen out-of-distribution (OOD) data in real-world scenarios. To address these issues, we propose a novel method for both ID classification and OOD detection (MIntOOD). We first introduce a weighted feature fusion network that models multimodal representations. This network dynamically learns the importance of each modality, adapting to multimodal contexts. To develop discriminative representations for both tasks, we synthesize pseudo-OOD data from convex combinations of ID data and engage in multimodal representation learning from both coarse-grained and fine-grained perspectives. The coarse-grained perspective focuses on distinguishing between ID and OOD binary classes, while the fine-grained perspective not only enhances the discrimination between different ID classes but also captures instance-level interactions between ID and OOD samples, promoting proximity among similar instances and separation from dissimilar ones. We establish baselines for three multimodal intent datasets and build an OOD benchmark. Extensive experiments on these datasets demonstrate that our method significantly improves OOD detection performance with a 3~10% increase in AUROC scores while achieving new state-of-the-art results in ID classification. Data and codes are available at https://github.com/thuiar/MIntOOD.

Figures

Figures reproduced from arXiv: 2412.12453 by the authors.

Figure 1
Figure 1. Examples of in-distribution and out-of-distribution samples for [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overall architecture of MIntOOD. It begins by generating pseudo-OOD samples through convex combinations of features extracted from ID samples. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. A comparison between different OOD detection methods. [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Confusion matrices for ID classes across the three datasets. [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: Distribution of OOD detection scores for ID and OOD data in the testing sets of the three datasets. [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark

    cs.CL 2025-04 conditional novelty 6.0 of 10

    MMLA combines 61K multimodal utterances across six semantic dimensions; even the best fine-tuned multimodal LLM reaches only about 69% accuracy, exposing current limits in cognitive-level language understanding.

  2. LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition

    cs.MM 2025-09 conditional novelty 5.0 of 10

    LGSRR uses LLM-generated semantic descriptions and rankings to improve multimodal intent recognition, reporting SOTA results on MIntRec2.0 and IEMOCAP-DA with gains around 0.5-1.3%.

  3. Deep Learning Approaches for Multimodal Intent Recognition: A Survey

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A survey of deep learning methods for intent recognition, tracing the field from unimodal text, audio, vision, and EEG approaches to multimodal fusion, alignment, knowledge-augmented, and multi-task models.

Reference graph

Works this paper leans on

75 extracted references · 69 canonical work pages · cited by 3 Pith papers

  1. [1]

    Mintrec: A new dataset for multimodal intent recognition,

    H. Zhang, H. Xu, X. Wang, Q. Zhou, S. Zhao, and J. Teng, “Mintrec: A new dataset for multimodal intent recognition,” in Proc. 30th ACM Int. Conf. on Multimedia , 2022, pp. 1688–1697

  2. [2]

    Multimodal activity detection for natural interaction with virtual human,

    K. Wang, S. Lian, H. Sang, W. Liu, Z. Liu, F. Shi, H. Deng, Z. Sun, and Z. Chen, “Multimodal activity detection for natural interaction with virtual human,” in Proc. 2023 IEEE Conf. Virtual Reality 3D User Interfaces Abstr. Workshops, 2023, pp. 671–672

  3. [3]

    Building multi-turn query interpreters for e-commercial chatbots with sparse-to-dense attentive modeling,

    Y . Fan, C. Wang, P. He, and Y . Hu, “Building multi-turn query interpreters for e-commercial chatbots with sparse-to-dense attentive modeling,” in Proc. 15th ACM Int. Conf. Web Search Data Mining , 2022, pp. 1577–1580

  4. [4]

    Intent based multimodal speech and gesture fusion for human-robot commu- nication in assembly situation,

    S. Paul, M. Sintek, V . K ¨epuska, M. Silaghi, and L. Robertson, “Intent based multimodal speech and gesture fusion for human-robot commu- nication in assembly situation,” in Proc. 2022 21st IEEE/CVF Int. Conf. Mach. Learn. Appl. , 2022, pp. 760–763

  5. [5]

    Towards a multimodal and context-aware framework for hu- man navigational intent inference,

    Z. Zhang, “Towards a multimodal and context-aware framework for hu- man navigational intent inference,” in Proc. 2020 Int. Conf. Multimodal Interaction, 2020, pp. 738–742

  6. [6]

    Object affordance based multimodal fusion for natural human-robot interaction,

    J. Mi, S. Tang, Z. Deng, M. G ¨orner, and J. Zhang, “Object affordance based multimodal fusion for natural human-robot interaction,” Cogn. Syst. Res., vol. 54, pp. 128–137, 2019

  7. [7]

    Multimodal transformer for unaligned multimodal language sequences,

    Y .-H. H. Tsai, S. Bai, P. P. Liang, J. Z. Kolter, L.-P. Morency, and R. Salakhutdinov, “Multimodal transformer for unaligned multimodal language sequences,” in Proc. 57th Assoc. Comput. Linguist. , 2019, pp. 6558–6569

  8. [8]

    Misa: Modality-invariant and-specific representations for multimodal sentiment analysis,

    D. Hazarika, R. Zimmermann, and S. Poria, “Misa: Modality-invariant and-specific representations for multimodal sentiment analysis,” in Proc. 28th ACM Int. Conf. on Multimedia , 2020, pp. 1122–1131

Show all 75 references
  1. [9]

    Integrating multimodal information in large pretrained transformers,

    W. Rahman, M. K. Hasan, S. Lee, A. Zadeh, C. Mao, L.-P. Morency, and E. Hoque, “Integrating multimodal information in large pretrained transformers,” in Proc. 58th Assoc. Comput. Linguist. , 2020, pp. 2359– 2369

  2. [10]

    Sdif-da: A shallow- to-deep interaction framework with data augmentation for multi-modal intent detection,

    S. Huang, L. Qin, B. Wang, G. Tu, and R. Xu, “Sdif-da: A shallow- to-deep interaction framework with data augmentation for multi-modal intent detection,” in Proc. IEEE/CVF Int. Conf. Acoust., Speech Signal Process., 2024, pp. 10 206–10 210

  3. [11]

    Token-level contrastive learning with modality-aware prompting for multimodal intent recognition,

    Q. Zhou, H. Xu, H. Li, H. Zhang, X. Zhang, Y . Wang, and K. Gao, “Token-level contrastive learning with modality-aware prompting for multimodal intent recognition,” in Proc. AAAI Conf. Artif. Intell. , 2024, pp. 17 114–17 122

  4. [12]

    Deep open intent classification with adaptive decision boundary,

    H. Zhang, H. Xu, and T.-E. Lin, “Deep open intent classification with adaptive decision boundary,” in Proc. AAAI Conf. Artif. Intell. , 2021, pp. 14 374–14 382

  5. [13]

    Learning discriminative representations and decision boundaries for open intent detection,

    H. Zhang, H. Xu, S. Zhao, and Q. Zhou, “Learning discriminative representations and decision boundaries for open intent detection,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 31, pp. 1611– 1623, 2023

  6. [14]

    An evaluation dataset for intent classification and out-of-scope prediction,

    S. Larson, A. Mahendran, J. J. Peper, C. Clarke, A. Lee, P. Hill, J. K. Kummerfeld, K. Leach, M. A. Laurenzano, L. Tang, and J. Mars, “An evaluation dataset for intent classification and out-of-scope prediction,” in Proc. 2019 Conf. Empir. Methods Nat. Lang. Process. and 9th I...

  7. [15]

    Generalized odin: Detecting out-of-distribution image without learning from out-of-distribution data,

    Y . C. Hsu, Y . Shen, H. Jin, and Z. Kira, “Generalized odin: Detecting out-of-distribution image without learning from out-of-distribution data,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2020, pp. 10 951–10 960

  8. [16]

    Energy-based out-of-distribution detection,

    W. Liu, X. Wang, J. Owens, and Y . Li, “Energy-based out-of-distribution detection,” in Proc. Adv. Neural Inf. Process. Syst. , 2020, pp. 21 464– 21 475

  9. [17]

    KNN-contrastive learning for out-of- domain intent classification,

    Y . Zhou, P. Liu, and X. Qiu, “KNN-contrastive learning for out-of- domain intent classification,” in Proc. 60th Assoc. Comput. Linguist. , 2022, pp. 5129–5141

  10. [18]

    Out-of-scope intent detection with self-supervision and discriminative training,

    L.-M. Zhan, H. Liang, B. Liu, L. Fan, X.-M. Wu, and A. Y . Lam, “Out-of-scope intent detection with self-supervision and discriminative training,” in Proc. 59th Assoc. Comput. Linguist. , 2021, pp. 3521–3532

  11. [19]

    Zero-shot out-of- distribution detection based on the pre-trained model clip,

    S. Esmaeilpour, B. Liu, E. Robertson, and L. Shu, “Zero-shot out-of- distribution detection based on the pre-trained model clip,” inProc. AAAI Conf. Artif. Intell. , 2022, pp. 6568–6576

  12. [20]

    Delving into out-of- distribution detection with vision-language representations,

    Y . Ming, Z. Cai, J. Gu, Y . Sun, W. Li, and Y . Li, “Delving into out-of- distribution detection with vision-language representations,” Proc. Adv. Neural Inf. Process. Syst. , pp. 35 087–35 102, 2022

  13. [21]

    A survey on spoken language understanding: Recent advances and new frontiers,

    L. Qin, T. Xie, W. Che, and T. Liu, “A survey on spoken language understanding: Recent advances and new frontiers,” in Proc. Int. Joint Conf. Artif. Intell. , 2021, pp. 4577–4584

  14. [22]

    Snips voice platform: an embedded spoken language understanding system for private-by-design voice interfaces,

    A. Coucke, A. Saade, A. Ball, T. Bluche, A. Caulier, D. Leroy, C. Doumouro, T. Gisselbrecht, F. Caltagirone, T. Lavril et al. , “Snips voice platform: an embedded spoken language understanding system for private-by-design voice interfaces,” arXiv preprint arXiv:1805.10190 , 2018

  15. [23]

    Efficient intent detection with dual sentence encoders,

    I. Casanueva, T. Tem ˇcinas, D. Gerz, M. Henderson, and I. Vuli ´c, “Efficient intent detection with dual sentence encoders,” in Proc. 2nd Workshop Nat. Lang. Process. for Conversational AI , 2020, pp. 38–45

  16. [24]

    Bert for joint intent classification and slot filling,

    Q. Chen, Z. Zhuo, and W. Wang, “Bert for joint intent classification and slot filling,” arXiv preprint arXiv:1902.10909 , 2019

  17. [25]

    Unified dialog model pre-training for task-oriented dialog understanding and generation,

    W. He, Y . Dai, M. Yang, J. Sun, F. Huang, L. Si, and Y . Li, “Unified dialog model pre-training for task-oriented dialog understanding and generation,” in Proc. 45th Int. ACM SIGIR Conf. Res. Develop. Inf. Retrieval, 2022, pp. 187–200

  18. [26]

    Discovering new intents via con- strained deep adaptive clustering with cluster refinement,

    T.-E. Lin, H. Xu, and H. Zhang, “Discovering new intents via con- strained deep adaptive clustering with cluster refinement,” in Proc. AAAI Conf. Artif. Intell. , 2020, pp. 8360–8367

  19. [27]

    Discovering new intents with deep aligned clustering,

    H. Zhang, H. Xu, T.-E. Lin, and R. Lyu, “Discovering new intents with deep aligned clustering,” in Proc. AAAI Conf. Artif. Intell. , 2021, pp. 14 365–14 373

  20. [28]

    A probabilistic framework for discov- ering new intents,

    Y . Zhou, G. Quan, and X. Qiu, “A probabilistic framework for discov- ering new intents,” in Proc. 61st Assoc. Comput. Linguist. , 2023, pp. 3771–3784

  21. [29]

    A clustering framework for unsupervised and semi-supervised new intent discovery,

    H. Zhang, H. Xu, X. Wang, F. Long, and K. Gao, “A clustering framework for unsupervised and semi-supervised new intent discovery,” IEEE Trans. Knowl. Data Eng. , vol. 36, pp. 5468–5481, 2023

  22. [30]

    Integrating text and image: Determining multimodal document intent 14 in Instagram posts,

    J. Kruk, J. Lubin, K. Sikka, X. Lin, D. Jurafsky, and A. Divakaran, “Integrating text and image: Determining multimodal document intent 14 in Instagram posts,” in Proc. 2019 Conf. Empir. Methods Nat. Lang. Process. and 9th Int. Joint Conf. Nat. Lang. Process. , 2019, pp. 4622– 4632

  23. [31]

    Multimodal marketing intent analysis for effective targeted advertising,

    L. Zhang, J. Shen, J. Zhang, J. Xu, Z. Li, Y . Yao, and L. Yu, “Multimodal marketing intent analysis for effective targeted advertising,” IEEE Trans. Multimedia, vol. 24, pp. 1830–1843, 2021

  24. [32]

    Emoint- trans: A multimodal transformer for identifying emotions and intents in social conversations,

    G. V . Singh, M. Firdaus, A. Ekbal, and P. Bhattacharyya, “Emoint- trans: A multimodal transformer for identifying emotions and intents in social conversations,” IEEE/ACM Trans. Audio, Speech, Lang. Process., vol. 31, pp. 290–300, 2022

  25. [33]

    Speech-text pre-training for spoken dialog understanding with explicit cross-modal alignment,

    T. Yu, H. Gao, T.-E. Lin, M. Yang, Y . Wu, W. Ma, C. Wang, F. Huang, and Y . Li, “Speech-text pre-training for spoken dialog understanding with explicit cross-modal alignment,” in Proc. 61st Assoc. Comput. Linguist., 2023, pp. 7900–7913

  26. [34]

    A deep multi-task model for dialogue act classification, intent detection and slot filling,

    M. Firdaus, H. Golchha, A. Ekbal, and P. Bhattacharyya, “A deep multi-task model for dialogue act classification, intent detection and slot filling,” Cogn. Comput., vol. 13, pp. 626–645, 2021

  27. [35]

    Switchboard: telephone speech corpus for research and development,

    J. Godfrey, E. Holliman, and J. McDaniel, “Switchboard: telephone speech corpus for research and development,” in Proc. 1992 IEEE/CVF Int. Conf. Acoust., Speech, Signal Process. , 1992, pp. 517–520

  28. [36]

    Dailydia- log: A manually labelled multi-turn dialogue dataset,

    Y . Li, H. Su, X. Shen, W. Li, Z. Cao, and S. Niu, “Dailydia- log: A manually labelled multi-turn dialogue dataset,” arXiv preprint arXiv:1710.03957, 2017

  29. [37]

    Towards emotion- aided multi-modal dialogue act classification,

    T. Saha, A. Patra, S. Saha, and P. Bhattacharyya, “Towards emotion- aided multi-modal dialogue act classification,” in Proc. 58th Assoc. Comput. Linguist., 2020, pp. 4361–4372

  30. [38]

    Meld: A multimodal multi-party dataset for emotion recognition in conversations,

    S. Poria, D. Hazarika, N. Majumder, G. Naik, E. Cambria, and R. Mihal- cea, “Meld: A multimodal multi-party dataset for emotion recognition in conversations,” in Proc. 57th Assoc. Comput. Linguist. , 2019, pp. 527–536

  31. [39]

    Iemocap: Interactive emotional dyadic motion capture database,

    C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,” Lang. Resour. Eval., vol. 42, pp. 335– 359, 2008

  32. [40]

    Ch-sims: A chinese multimodal sentiment analysis dataset with fine- grained annotation of modality,

    W. Yu, H. Xu, F. Meng, Y . Zhu, Y . Ma, J. Wu, J. Zou, and K. Yang, “Ch-sims: A chinese multimodal sentiment analysis dataset with fine- grained annotation of modality,” in Proc. 58th Assoc. Comput. Linguist., 2020, pp. 3718–3727

  33. [41]

    Mosi: multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos,

    A. Zadeh, R. Zellers, E. Pincus, and L.-P. Morency, “Mosi: multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos,” arXiv preprint arXiv:1606.06259 , 2016

  34. [42]

    Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,

    A. B. Zadeh, P. P. Liang, S. Poria, E. Cambria, and L.-P. Morency, “Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,” in Proc. 56th Assoc. Comput. Linguist., 2018, pp. 2236–2246

  35. [43]

    MIntrec2.0: A large-scale benchmark dataset for multi- modal intent recognition and out-of-scope detection in conversations,

    H. Zhang, X. Wang, H. Xu, Q. Zhou, J. Su, jinyue Zhao, W. Li, Y . Chen, and K. Gao, “MIntrec2.0: A large-scale benchmark dataset for multi- modal intent recognition and out-of-scope detection in conversations,” in Proc. 12th Int. Conf. Learn. Represent. , 2024

  36. [44]

    Tensor fusion network for multimodal sentiment analysis,

    A. Zadeh, M. Chen, S. Poria, E. Cambria, and L.-P. Morency, “Tensor fusion network for multimodal sentiment analysis,” in Proc. 2017 Conf. Empir. Methods Nat. Lang. Process. , 2017, pp. 1103–1114

  37. [45]

    Efficient low-rank multimodal fusion with modality- specific factors,

    Z. Liu, Y . Shen, V . B. Lakshminarasimhan, P. P. Liang, A. Bagher Zadeh, and L.-P. Morency, “Efficient low-rank multimodal fusion with modality- specific factors,” inProc. 56th Assoc. Comput. Linguist., 2018, pp. 2247– 2256

  38. [46]

    Memory fusion network for multi-view sequential learning,

    A. Zadeh, P. P. Liang, N. Mazumder, S. Poria, E. Cambria, and L.-P. Morency, “Memory fusion network for multi-view sequential learning,” in Proc. AAAI Conf. Artif. Intell. , 2018, pp. 5634–5641

  39. [47]

    Improving multimodal fusion with hierarchical mutual information maximization for multimodal sentiment analysis,

    W. Han, H. Chen, and S. Poria, “Improving multimodal fusion with hierarchical mutual information maximization for multimodal sentiment analysis,” in Proc. 2021 Conf. Empir. Methods Nat. Lang. Process., 2021, pp. 9180–9192

  40. [48]

    Unimse: Towards unified multimodal sentiment analysis and emotion recognition,

    G. Hu, T.-E. Lin, Y . Zhao, G. Lu, Y . Wu, and Y . Li, “Unimse: Towards unified multimodal sentiment analysis and emotion recognition,” inProc. 2022 Conf. Empir. Methods Nat. Lang. Process. , 2022, pp. 7837–7851

  41. [49]

    Exploring the limits of transfer learning with a unified text-to-text transformer,

    C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res. , vol. 21, pp. 1–67, 2020

  42. [50]

    Out-of-domain detection for natural language understanding in dialog systems,

    Y . Zheng, G. Chen, and M. Huang, “Out-of-domain detection for natural language understanding in dialog systems,” IEEE/ACM Trans. Audio, Speech, Lang. Process., vol. 28, pp. 1198–1209, 2020

  43. [51]

    Improving unsupervised out- of-domain detection through pseudo labeling and learning,

    B. Lee, J. Kim, J. Park, and K.-A. Sohn, “Improving unsupervised out- of-domain detection through pseudo labeling and learning,” in Proc. Findings of the Assoc. Comput. Linguist.: EACL , 2023, pp. 1031–1041

  44. [52]

    Enhancing the generalization for intent classification and out-of-domain detection in slu,

    Y . Shen, Y .-C. Hsu, A. Ray, and H. Jin, “Enhancing the generalization for intent classification and out-of-domain detection in slu,” in Proc. 59th Assoc. Comput. Linguist. , 2021, pp. 2443–2453

  45. [53]

    A closer look at few-shot out-of-distribution intent detection,

    L.-M. Zhan, H. Liang, L. Fan, X.-M. Wu, and A. Y . Lam, “A closer look at few-shot out-of-distribution intent detection,” in Proc. 29th Int. Conf. Comput. Linguist. , 2022, pp. 451–460

  46. [54]

    A simple unified framework for detecting out-of-distribution samples and adversarial attacks,

    K. Lee, K. Lee, H. Lee, and J. Shin, “A simple unified framework for detecting out-of-distribution samples and adversarial attacks,” in Proc. Adv. Neural Inf. Process. Syst. , 2018, pp. 7166–7177

  47. [55]

    Out-of-distribution detection with subspace techniques and probabilistic modeling of features,

    I. Ndiour, N. Ahuja, and O. Tickoo, “Out-of-distribution detection with subspace techniques and probabilistic modeling of features,” arXiv preprint arXiv:2012.04250, 2020

  48. [56]

    Vim: Out-of-distribution with virtual-logit matching,

    H. Wang, Z. Li, L. Feng, and W. Zhang, “Vim: Out-of-distribution with virtual-logit matching,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2022, pp. 4921–4930

  49. [57]

    Swin transformer: Hierarchical vision transformer using shifted windows,

    Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , 2021, pp. 9992– 10 002

  50. [58]

    librosa: Audio and music signal analysis in python,

    B. McFee, C. Raffel, D. Liang, D. P. Ellis, M. McVicar, E. Battenberg, and O. Nieto, “librosa: Audio and music signal analysis in python,” in Proc. 14th Python in Sci. Conf. , 2015, pp. 18–25

  51. [59]

    Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

    S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “Wavlm: Large-scale self-supervised pre- training for full stack speech processing,” IEEE J. Sel. Top. Signal Process., vol. 16, pp. 1505–1518, 2022

  52. [60]

    Energy- based unknown intent detection with data manipulation,

    Y . Ouyang, J. Ye, Y . Chen, X. Dai, S. Huang, and J. Chen, “Energy- based unknown intent detection with data manipulation,” in Findings of the Assoc. Comput. Linguist.: ACL-IJCNLP , 2021, pp. 2852–2861

  53. [61]

    Learning to classify open intent via soft labeling and manifold mixup,

    Z. Cheng, Z. Jiang, Y . Yin, C. Wang, and Q. Gu, “Learning to classify open intent via soft labeling and manifold mixup,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 30, pp. 635–645, 2022

  54. [62]

    Bert: Pre-training of deep bidirectional transformers for language understanding,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proc. 2019 Conf. North Am. Chapter Assoc. Comput. Linguist.: Human Lang. Technol., 2019, pp. 4171–4186

  55. [63]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. Adv. Neural Inf. Process. Syst. , 2017, pp. 5998–6008

  56. [64]

    Generalized out-of-distribution detection: A survey,

    J. Yang, K. Zhou, Y . Li, and Z. Liu, “Generalized out-of-distribution detection: A survey,” Int. J. Comput. Vis. , vol. 132, pp. 5635–5662, 2024

  57. [65]

    App: Adaptive prototypical pseudo-labeling for few-shot ood detection,

    P. Wang, K. He, Y . Mou, X. Song, Y . Wu, J. Wang, Y . Xian, X. Cai, and W. Xu, “App: Adaptive prototypical pseudo-labeling for few-shot ood detection,” in Findings of the Assoc. Comput. Linguist.: EMNLP , 2023, pp. 3926–3939

  58. [66]

    Low-shot learning with imprinted weights,

    H. Qi, M. Brown, and D. G. Lowe, “Low-shot learning with imprinted weights,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2018, pp. 5822–5830

  59. [67]

    Simcse: Simple contrastive learning of sentence embeddings,

    T. Gao, X. Yao, and D. Chen, “Simcse: Simple contrastive learning of sentence embeddings,” in Proc. 2021 Conf. Empir. Methods Nat. Lang. Process., 2021, pp. 6894–6910

  60. [68]

    Dialogue act modeling for automatic tagging and recognition of conversational speech,

    A. Stolcke, K. Ries, N. Coccaro, E. Shriberg, R. Bates, D. Jurafsky, P. Taylor, R. Martin, C. V . Ess-Dykema, and M. Meteer, “Dialogue act modeling for automatic tagging and recognition of conversational speech,” Comput. Linguist., vol. 26, pp. 339–373, 2000

  61. [69]

    A baseline for detecting misclassified and out-of-distribution examples in neural networks,

    D. Hendrycks and K. Gimpel, “A baseline for detecting misclassified and out-of-distribution examples in neural networks,” in Proc. 5th Int. Conf. Learn. Represent. , 2017

  62. [70]

    Enhancing the reliability of out-of- distribution image detection in neural networks,

    S. Liang, Y . Li, and R. Srikant, “Enhancing the reliability of out-of- distribution image detection in neural networks,” in Proc. 6th Int. Conf. Learn. Represent., 2018

  63. [71]

    Representation learning with contrastive predictive coding,

    A. v. d. Oord, Y . Li, and O. Vinyals, “Representation learning with contrastive predictive coding,” arXiv preprint arXiv:1807.03748 , 2018

  64. [72]

    Out-of-domain detection based on generative adversarial network,

    S. Ryu, S. Koo, H. Yu, and G. G. Lee, “Out-of-domain detection based on generative adversarial network,” in Proc. 2018 Conf. Empir. Methods Nat. Lang. Process. , 2018, pp. 714–718

  65. [73]

    Transformers: State-of- the-art natural language processing,

    T. Wolf, L. Debut, V . Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowiczet al., “Transformers: State-of- the-art natural language processing,” inProc. 2020 Conf. Empir. Methods Nat. Lang. Process.: Sys. Demonstrations , 2020, pp. 38–45

  66. [74]

    Decoupled weight decay regularization,

    I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in Proc. 7th Int. Conf. Learn. Represent. , 2019

  67. [75]

    Out-of-distribution detection using union of 1- dimensional subspaces,

    A. Zaeemzadeh, N. Bisagno, Z. Sambugaro, N. Conci, N. Rah- navard, and M. Shah, “Out-of-distribution detection using union of 1- dimensional subspaces,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2021, pp. 9452–9461. 15 Hanlei Zhang received the B.S. degree from t...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.