Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Deep Learning Approaches for Multimodal Intent Recognition: A Survey

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This survey traces intent recognition from text-only models to multimodal deep learning, organizing the field around ten benchmark datasets and a four-paradigm taxonomy.

desk verdict A useful organizing survey whose 'first systematic review' claim is not backed by any disclosed methodology, and whose curation errors make the synthesis hard to check. read the letter →

arxiv 2507.22934 v1 pith:7AIZDHTC submitted 2025-07-24 cs.CL cs.AI

classification cs.CLcs.AI
keywords multimodalintentrecognitionlearningdeepsurveybenchmarkdatasetstaxonomylargelanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to establish a structured map of deep learning for intent recognition, tracing the field from text-only models through vision, audio, and EEG to multimodal systems that fuse several signals at once. Its central claim is that multimodal intent recognition is the natural next stage because single modalities fail under noise, ambiguity, and missing context. To make the field tractable, the paper organizes methods around a three-stage pipeline and a four-paradigm taxonomy, and it compiles the widely used datasets and evaluation metrics into a single reference. A reader would care because the survey gives researchers a common vocabulary, a benchmark baseline, and an explicit list of open problems—ambiguity, multi-intent utterances, evolving dialogue intent, modality asynchrony, out-of-domain inputs, long-tail labels, cross-lingual gaps, and continuous reasoning in dynamic environments.

What carries the argument

The load-bearing organizing device is the three-stage multimodal intent recognition pipeline—Feature Extraction, Multimodal Representation Learning, and Intent Classification—with representation learning treated as the core stage. Within that stage, the paper's central taxonomy divides methods into four paradigms: Fusion Methods (feature-, decision-, and hybrid-level), Alignment and Disentanglement Methods (contrastive learning, cross-modal attention, disentangled encoders), Knowledge-Augmented Methods (LLM-based and retrieval-based), and Multi-Task Coordination Methods (joint optimization of intent with related tasks such as emotion recognition). This machinery is what lets the survey place dozens of separate model papers into a single comparative structure.

What would settle it

A comprehensive, reproducible literature search over the same period that finds either an earlier survey covering the same unimodal-to-multimodal trajectory, or a substantial cluster of multimodal intent recognition methods that cannot be placed into any of the four paradigms (fusion, alignment and disentanglement, knowledge-augmented, multi-task coordination), would refute the paper's central claims; a concrete version would count what fraction of a random sample of 2019–2025 multimodal intent recognition papers falls outside the four paradigms.

Watch

Extended reading notes

Core claim

In the paper's own framing, the core discovery is that intent recognition has undergone a coherent evolution that can be systematically reviewed: initial rule- and feature-based text methods gave way to deep learning, and increasingly to Transformer- and LLM-based models, while the field simultaneously expanded from text to vision, audio, EEG, and multimodal combinations. The authors claim to present the first systematic review covering this whole trajectory, and they support the claim with a catalog of ten datasets, a coarse-grained intent taxonomy (Emotion and Attitude, Goal Achievement, Information and Declaration), and a classification of multimodal methods into fusion, alignment and disentanglement, knowledge-augmented, and multi-task coordination paradigms. They further assemble evaluation metrics, including Accuracy, Precision, Recall, F1, and specialized out-of-scope measures, and identify application areas from human-computer interaction to automotive systems and sports.

Load-bearing premise

The survey's map of the field relies on a hand-picked collection of datasets and methods chosen without a stated systematic search or inclusion criteria, so the proposed pipeline and taxonomy hold only if that selection is representative.

Editorial extensions

If this is right

  • New multimodal intent recognition systems can be described and compared against a common pipeline, so performance numbers from different papers become easier to relate to one another.
  • The dataset catalog of ten benchmarks gives researchers a ready-made evaluation foundation for unimodal and multimodal settings, including out-of-scope detection benchmarks.
  • The four-paradigm taxonomy suggests that future work will increasingly combine paradigms—for example, knowledge augmentation used inside a contrastive alignment framework—rather than staying within one of them.
  • The eight named challenges act as a research agenda: work on ambiguity, multi-intent structure, dialogue-level intent evolution, modality asynchrony, OOD detection, long-tail labels, cross-lingual generalization, and continuous reasoning are all identified as open problems.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The survey's implicit emphasis on text as the anchor modality suggests that robustness to missing modalities—models that must still work when video or audio is absent—could become the next standard evaluation criterion; the paper lists this issue but does not develop a benchmark protocol for it.
  • A testable consequence of the four-paradigm taxonomy is coverage: applying the taxonomy to a broader, systematically sampled set of 2019–2025 papers would reveal whether any substantial family of multimodal intent methods falls outside the four paradigms.
  • The inclusion of EEG and eye-tracking datasets points to cognitive-signal fusion as an underexplored frontier, where intent is inferred from neural and gaze data rather than spoken or typed language; this direction is visible in the survey's dataset table but not developed as a research program.
  • The recurring use of large language models for label description, knowledge extraction, and reasoning suggests that the field may converge on LLM-generated pseudo-labels or knowledge as a standard data-augmentation step, an implication the survey leaves implicit.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This survey reviews deep learning methods for intent recognition, tracing the field from unimodal text/vision/audio/EEG approaches to multimodal intent recognition (MIR). It proposes a three-stage processing pipeline (feature extraction, multimodal representation learning, intent classification) and organizes MIR methods into four paradigms: fusion, alignment and disentanglement, knowledge-augmented, and multi-task coordination. The paper compiles benchmark datasets, representative methods, evaluation metrics, applications, and challenges, and it claims to be the first systematic review of this trajectory. The survey is curated rather than derived: there are no fitted parameters or formal derivations, and the central contributions are the taxonomy, the resource inventory, and the field-level synthesis.

Significance. If the inventory and taxonomy are reliable, the paper is a useful reference for researchers entering multimodal intent recognition, especially because it covers many 2024–2025 works and organizes them into an explicit taxonomy. The evaluation-metrics section, with equations for Macro F1, Weighted F1, F1-IS/F1-OOS, and EER, is a practical contribution. The application and challenge sections are broad and well organized, and the paper explicitly names open problems such as modal asynchrony, long-tail distributions, and cross-lingual generalization. The main caveat is that the 'systematic review' claim is not currently supported by a disclosed methodology, and the dataset/method inventory contains internal inconsistencies. These issues are fixable, but they are load-bearing for the paper's comparative and synthetic conclusions.

major comments (3)
  1. [§1 (first contribution bullet), §2.1, §4.5] The first contribution bullet claims 'the first systematic review that traces the development of intent recognition from early unimodal approaches to modern multimodal techniques,' but the manuscript provides no search protocol, query strings, inclusion/exclusion criteria, screening procedure, or comparison with existing surveys of similar scope such as [81], [143], and [3]. Without this methodology, the dataset and method sample in Tables 2 and 4 cannot be checked for representativeness, and the field-level generalizations in §4.5 — e.g., that fusion methods 'effectively combine text, audio, and visual signals' or that knowledge-augmented methods are 'limited in scalability' — rest on an unverified curated sample. Please either add a transparent systematic methodology or revise the claim to 'structured survey' and temper the corresponding generalizations.
  2. [§2.1, Table 2, §3.2, Table 4] Several internal consistency errors affect the resource inventory. MultiWOZ is described as containing '10,438 samples (8,438 conversations)'; in the original corpus 10,438 is the dialogue count and 8,438 is the training split, so the sample/dialogue counts are inverted. MDID is used as a benchmark in §3.2 (Tang et al. [124]) but is absent from Table 2. Methods discussed in the text — MIntOOD [164], KDSL [10], and MMSAIR [116] — are missing from Table 4. Please correct the MultiWOZ statistics and align the tables with the narrative so that the 'standardized foundation' claim is credible.
  3. [§4.1–§4.4] The four methodological paradigms are presented as a partition of MIR research, but membership rules are not given, and several methods fall into multiple categories: CaVIR appears under fusion, alignment, and knowledge-augmented methods; MGC appears under fusion and alignment; A-MESS appears under alignment and knowledge-augmented methods. It is not clear whether these are mutually exclusive classes or complementary aspects of a design space. Please state the classification criterion explicitly, and either assign each method to one primary paradigm or state that the categories are overlapping facets.
minor comments (5)
  1. [References] References [76] and [77] are duplicate entries for the same MBCFNet paper, and references [134] and [135] are duplicate entries for the same MGC paper; please merge the duplicates and update the in-text citations accordingly.
  2. [§5] The metrics list introduces Weighted Precision (WP) and Pearson Correlation (Corr), but neither is defined or used later; please provide formulas or remove them from the list.
  3. [Table 3] The Dataset column uses the abbreviation 'Comm.' in several rows without explanation; please define this abbreviation or replace it with the actual dataset names.
  4. [§3.3] In the discussion of MuProCL, the text says 'SLURP and MintRec'; the dataset name should be 'MIntRec'.
  5. [§2.1] Item (1) covers five text datasets while items (2)–(10) each cover a single dataset; renumbering or restructuring would make the dataset list easier to navigate.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey makes no predictive derivation, has no fitted parameters, and its taxonomy is explicitly attributed to existing datasets.

full rationale

This manuscript is a literature survey rather than a derivation: there are no fitted parameters, no equations that reduce to inputs, and no claimed first-principles result. The only construction in the paper is the coarse-grained intent taxonomy in Section 2.1, which is explicitly described as 'Inspired by the classification approach of the MIntRec dataset' and by 'distribution patterns of intentions across domains'; it is presented as an organizational choice, not as a prediction or theorem, so it does not constitute a circular step. The 'first systematic review' contribution in Section 1 is a novelty claim unsupported by a disclosed search protocol, but lack of evidence for a systematic methodology is a correctness/representativeness concern, not a circularity concern. No load-bearing argument is justified by a self-citation: the MIntRec and MIntRec2.0 references are not authored by the present authors, and the survey's claims do not reduce to any cited prior work. Duplicate references (e.g., [76]/[77] and [134]/[135]) and dataset description issues are curation errors, not circular reasoning. Accordingly, the circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

This is a survey, so there are no fitted parameters, no derivation, and no newly postulated entities. The load-bearing assumptions are about the representativeness and descriptive adequacy of the paper's taxonomies and selected literature.

assumptions (3)
  • domain assumption The chosen coarse-grained taxonomy (Emotion and Attitude, Goal Achievement, Information and Declaration) is a faithful abstraction of intent categories across domains.
    Introduced in Section 2.1, based on MIntRec's classification approach; the survey's organization of datasets and methods depends on it.
  • domain assumption The selected 10 datasets and representative methods are representative of the field.
    No systematic inclusion criteria are stated; Table 2 and Sections 3 and 4 use a curated list to support general claims about the field.
  • ad hoc to paper The three-stage pipeline and four-paradigm taxonomy (fusion, alignment and disentanglement, knowledge-augmented, multi-task coordination) faithfully partition multimodal intent recognition research.
    This is the paper's own organizing scheme, introduced in Section 4, and is not derived from a prior established taxonomy.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Deep Learning Approaches for Multimodal Intent Recognition: A Survey." pith.science (2026). https://pith.science/paper/7AIZDHTC

@misc{pith2026250722934,
  author       = {Pith},
  title        = {Pith review of: Deep Learning Approaches for Multimodal Intent Recognition: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7AIZDHTC}},
  note         = {Machine review of arXiv:2507.22934}
}
read the original abstract

Intent recognition aims to identify users' underlying intentions, traditionally focusing on text in natural language processing. With growing demands for natural human-computer interaction, the field has evolved through deep learning and multimodal approaches, incorporating data from audio, vision, and physiological signals. Recently, the introduction of Transformer-based models has led to notable breakthroughs in this domain. This article surveys deep learning methods for intent recognition, covering the shift from unimodal to multimodal techniques, relevant datasets, methodologies, applications, and current challenges. It provides researchers with insights into the latest developments in multimodal intent recognition (MIR) and directions for future research.

Figures

Figures reproduced from arXiv: 2507.22934 by the authors.

Figure 1
Figure 1. Representative Methods and Datasets for Intent Recognition. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Organizational Structure of the Survey. • We present the first systematic review that traces the development of intent recognition from early unimodal approaches to modern multimodal techniques, offering a structured and comparative perspective on the evolution of this field. • We collate and analyze benchmark datasets and evaluation metrics, covering both unimodal and multimodal settings, to offer researchers a sta… view at source ↗
Figure 3
Figure 3. Machine Learning for Text Intent Recognition. [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: The CNN/RNN-Based Model for Text Intent Recognition. [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: The BERT Model for Text Intent Recognition. [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]
Figure 6
Figure 6. Figure 6: The LLM-Based Models for Text Intent Recognition. (a) Leverages RAG (Retrieval-Augmented Generation) to concatenate [PITH_FULL_IMAGE:figures/full_fig_p012_6.png]
Figure 7
Figure 7. Figure 7: The Classification Paradigm and Prototype-Based Paradigm for Vision Intent Recognition. [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]
Figure 8
Figure 8. Figure 8: The Three Frameworks for Audio Intent Recognition (AIR). [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]
Figure 9
Figure 9. Figure 9: A Deep Learning Pipeline for Multimodal Intent Recognition. [PITH_FULL_IMAGE:figures/full_fig_p018_9.png]
Figure 10
Figure 10. Figure 10: The Basic Modal Fusion Methods of MIR. 4.1.1 Feature-Level Fusion. Multi-modal feature fusion methods based on attention mechanisms have demonstrated significant advantages in intent recognition tasks, effectively capturing cross-modal semantic associations, and have …
Figure 11
Figure 11. Figure 11: Single-Task Learning and Multi-Task Learning in MIR. [PITH_FULL_IMAGE:figures/full_fig_p021_11.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adaptive Modality Reliability Diagnosis and Restoration for Robust Multimodal Intent Recognition

    cs.MM 2026-08 conditional novelty 6.0 of 10

    A closed-loop framework that learns modality reliability via self-supervised corruption, restores unreliable modalities from reliable ones, re-checks reliability, and fuses with precision weights.

Reference graph

Works this paper leans on

178 extracted references · 57 canonical work pages · cited by 1 Pith paper

  1. [81]

    Jiao Liu, Yanling Li, and Min Lin. 2019. Review of Intent Detection Methods in the Human-Machine Dialogue System.Journal of Physics: Conference Series 1267, 1 (2019), 012059. doi:10.1088/1742-6596/1267/1/012059

  2. [143]

    Hua Xu, Hanlei Zhang, and Ting-En Lin. 2023. Intent Recognition for Human-Machine Interactions . Springer

  3. [76]

    Zhongjie Li, Gaoyan Zhang, Shogo Okada, Longbiao Wang, Bin Zhao, and Jianwu Dang. 2024. MBCFNet: A multimodal brain–computer fusion network for human intention recognition. Knowledge-Based Systems 296 (2024), 111826

  4. [77]

    Zhongjie Li, Gaoyan Zhang, Shogo Okada, Longbiao Wang, Bin Zhao, and Jianwu Dang. 2024. MBCFNet: A Multimodal Brain–Computer Fusion Network for human intention recognition. Knowledge-Based Systems 296 (2024), 111826. doi:10.1016/j.knosys.2024.111826

  5. [134]

    Mengsheng Wang, Lun Xie, Chiqin Li, Xinheng Wang, Minglong Sun, and Ziyang Liu. 2025. MGC: A modal mapping coupling and gate-driven contrastive learning approach for multimodal intent recognition. Expert Systems with Applications 281 (2025), 127631

  6. [135]

    Mengsheng Wang, Lun Xie, Chiqin Li, Xinheng Wang, Minglong Sun, and Ziyang Liu. 2025. MGC: A Modal Mapping Coupling and Gate-Driven Contrastive Learning Approach for Multimodal Intent Recognition. Expert Systems with Applications 281 (2025), 127631. doi:10.1016/j.eswa.2025. 127631

  7. [3]

    Jesse Atuhurra, Hidetaka Kamigaito, Taro Watanabe, and Eric Nichols. 2024. Domain Adaptation in Intent Classification Systems: A Review. arXiv preprint arXiv:2404.14415 (2024)

  8. [124]

    Yin Tang, Jiankai Li, Hongyu Yang, Xuan Dong, Lifeng Fan, and Weixin Li. 2025. Multi-Grained Compositional Visual Clue Learning for Image Intent Recognition. arXiv preprint arXiv:2504.18201 (2025)

  9. [164]

    Hanlei Zhang, Qianrui Zhou, Hua Xu, Jianhua Su, Roberto Evans, and Kai Gao. 2024. Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding. arXiv preprint arXiv:2412.12453 (2024)

  10. [10]

    Bin Chen, Yu Zhang, Hongfei Ye, Yizi Huang, and Hongyang Chen. 2025. Knowledge-Decoupled Synergetic Learning: An MLLM based Collaborative Approach to Few-shot Multimodal Dialogue Intention Recognition. In Companion Proceedings of the ACM on Web Conference 2025 . 3044–3048

  11. [116]

    Yuanchen Shi, Biao Ma, and Fang Kong. 2024. Impact of Stickers on Multimodal Chat Sentiment Analysis and Intent Recognition: A New Task, Dataset and Baseline. arXiv preprint arXiv:2405.08427 (2024)

Show all 178 references
  1. [1]

    Waheed Ahmed Abro, Guilin Qi, Huan Gao, Muhammad Asif Khan, and Zafar Ali. 2019. Multi-turn intent determination for goal-oriented dialogue systems. In International Joint Conference on Neural Networks (IJCNN) . IEEE, 1–8

  2. [2]

    Allen and C.Raymond Perrault

    James F. Allen and C.Raymond Perrault. 1980. Analyzing intention in utterances. Artificial Intelligence 15, 3 (1980), 143–178

  3. [4]

    Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information Processing Systems 33 (2020), 12449–12460

  4. [5]

    Emanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski, and Verena Rieser. 2020. SLURP: A spoken language understanding resource package. In 2020 Conference on Empirical Methods in Natural Language Processing . 7252–7262

  5. [6]

    Holger Berndt, Jorg Emmert, and Klaus Dietmayer. 2008. Continuous driver intention recognition with hidden markov models. In 2008 11th International IEEE Conference on Intelligent Transportation Systems . 1189–1194

  6. [7]

    Leo Breiman. 2001. Random forests. Machine Learning 45 (2001), 5–32

  7. [8]

    Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gašić. 2018. MultiWOZ - A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling. In Proceedings of the 2018 Conference on Empirical Metho...

  8. [9]

    Iñigo Casanueva, Tadas Temčinas, Daniela Gerz, Matthew Henderson, and Ivan Vulić. 2020. Efficient intent detection with dual sentence encoders. arXiv preprint arXiv:2003.04807 (2020)

  9. [11]

    Qian Chen, Zhu Zhuo, and Wen Wang. 2019. Bert for joint intent classification and slot filling. arXiv preprint arXiv:1902.10909 (2019)

  10. [12]

    Weitong Chen, Sen Wang, Xiang Zhang, Lina Yao, Lin Yue, Buyue Qian, and Xue Li. 2018. EEG-based motion intention recognition via multi-task RNNs. In Proceedings of the 2018 SIAM International Conference on Data Mining . SIAM, 279–287

  11. [13]

    Yongjun Chen, Zhiwei Liu, Jia Li, Julian McAuley, and Caiming Xiong. 2022. Intent contrastive learning for sequential recommendation. In Proceedings of the ACM Web Conference 2022 . 2172–2182

  12. [14]

    Yuan-Ping Chen, Ryan Price, and Srinivas Bangalore. 2018. Spoken language understanding without speech recognition. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6189–6193

  13. [15]

    Zhanpeng Chen, Zhihong Zhu, Xianwei Zhuang, Zhiqi Huang, and Yuexian Zou. 2024. Dual-oriented Disentangled Network with Counterfactual Intervention for Multimodal Intent Detection. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing . 17554– 17567

  14. [16]

    Xuxin Cheng, Zhihong Zhu, Hongxiang Li, Yaowei Li, Xianwei Zhuang, and Yuexian Zou. 2024. Towards multi-intent spoken language understanding via hierarchical attention and optimal transport. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 17844–17852

  15. [17]

    Fabrice Colas and Pavel Brazdil. 2006. Comparison of SVM and some older classification algorithms in text classification tasks. In IFIP International Conference on Artificial Intelligence in Theory and Practice . 169–178. 28 J. Zhao et al

  16. [18]

    Daniele Comi, Dimitrios Christofidellis, Pier Piazza, and Matteo Manica. 2023. Zero-Shot-BERT-adapters: A Zero-Shot Pipeline for Unknown Intent Detection. In Findings of the Association for Computational Linguistics: EMNLP 2023 . 650–663. doi:10.18653/v1/2023.findings-emnlp.47

  17. [19]

    Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Édouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised Cross-lingual Representation Learning at Scale. In Proceedings of the 58th Annual Meetin...

  18. [20]

    Corinna Cortes and Vladimir Vapnik. 1995. Support-vector networks. Machine Learning 20 (1995), 273–297

  19. [21]

    Alice Coucke, Alaa Saade, Adrien Ball, Théodore Bluche, Alexandre Caulier, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, et al. 2018. Snips voice platform: an embedded spoken language understanding system for private-by-design voi...

  20. [22]

    Thierry Desot, François Portet, and Michel Vacher. 2019. Towards end-to-end spoken intent recognition in smart home. In 2019 International Conference on Speech Technology and Human-Computer Dialogue (SpeD) . IEEE, 1–8

  21. [23]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human...

  22. [24]

    Pranay Dighe, Prateeth Nayak, Oggi Rudovic, Erik Marchi, Xiaochuan Niu, and Ahmed Tewfik. 2023. Audio-to-intent using acoustic-textual subword representations from end-to-end asr. InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICAS...

  23. [25]

    Qian Dong, Yuezhou Dong, Ke Qin, Guiduo Duan, and Tao He. 2025. Unbiased Multimodal Audio-to-Intent Recognition. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 1–5

  24. [26]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, G Heigold, S Gelly, et al. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Confere...

  25. [27]

    Veera Raghavendra Elluru, Devang Kulshreshtha, Rohit Paturi, Sravan Bodapati, and Srikanth Ronanki. 2023. Generalized zero-shot audio-to-intent classification. In 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 1–8

  26. [28]

    Mihail Eric, Rahul Goel, Shachi Paul, Abhishek Sethi, Sanchit Agarwal, Shuyang Gao, Adarsh Kumar, Anuj Goyal, Peter Ku, and Dilek Hakkani-Tur

  27. [29]

    Jianwu Fang, Fan Wang, Jianru Xue, and Tat-Seng Chua. 2024. Behavioral Intention Prediction in Driving Scenes: A Survey. IEEE Transactions on Intelligent Transportation Systems 25, 8 (2024), 8334–8355

  28. [30]

    Fatema Tuj Johora Faria, Mukaffi Bin Moin, Md Mahfuzur Rahman, Md Morshed Alam Shanto, Asif Iftekher Fahim, and Md Moinul Hoque. 2025. Uddessho: An extensive benchmark dataset for multimodal author intent classification in low-resource bangla language. InInternational Conferen...

  29. [31]

    Zihao Feng, Xiaoxue Wang, Ziwei Bai, Donghang Su, Bowen Wu, Qun Yu, and Baoxun Wang. 2025. Improving Generalization in Intent Detection: GRPO with Reward-Based Curriculum Sampling. arXiv preprint arXiv:2504.13592 (2025)

  30. [32]

    George Forman, Hila Nachlieli, and Renato Keshet. 2015. Clustering by Intent: A Semi-Supervised Method to Discover Relevant Clusters Incrementally. In Machine Learning and Knowledge Discovery in Databases . 20–36

  31. [33]

    Yoav Freund, Robert E Schapire, et al. 1996. Experiments with a new boosting algorithm. In icml, Vol. 96. 148–156

  32. [34]

    Tianhong Gao, Genhang Shen, Yuxuan Wu, Zunlei Feng, Jinshan Zhang, and Sheng Zhou. 2025. EcomMIR: Towards Intelligent Multimodal Intent Recognition in E-Commerce Dialogue Systems. In Companion Proceedings of the ACM on Web Conference 2025 . 3049–3052

  33. [35]

    Alexander Genkin, David D Lewis, and David Madigan. 2007. Large-scale Bayesian logistic regression for text categorization. Technometrics 49, 3 (2007), 291–304

  34. [36]

    Chih-Wen Goo, Guang Gao, Yun-Kai Hsu, Chih-Li Huo, Tsung-Chieh Chen, Keng-Wei Hsu, and Yun-Nung Chen. 2018. Slot-Gated Modeling for Joint Slot Filling and Intent Prediction. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computationa...

  35. [37]

    Dilek Hakkani-Tür, Gökhan Tür, Asli Celikyilmaz, Yun-Nung Chen, Jianfeng Gao, Li Deng, and Ye-Yi Wang. 2016. Multi-domain joint semantic frame parsing using bi-directional rnn-lstm.. In Interspeech. 715–719

  36. [38]

    Mohammad Mehedi Hassan, Stephen Karungaru, and Kenji Terada. 2024. Robotics Perception: Intention Recognition to Determine the Handball Occurrence during a Football or Soccer Match. AI 5, 2 (2024), 602–617

  37. [39]

    Gaole He, Nilay Aishwarya, and Ujwal Gadiraju. 2025. Is Conversational XAI All You Need? Human-AI Decision Making with a Conversational XAI Assistant. In Proceedings of the 30th International Conference on Intelligent User Interfaces (IUI ’25) . New York, NY, USA, 907–924. doi...

  38. [40]

    Charles T Hemphill, John J Godfrey, and George R Doddington. 1990. The ATIS spoken language systems pilot corpus. In Speech and Natural Language: Proceedings of a Workshop Held at Hidden Valley, Pennsylvania, June 24-27, 1990

  39. [41]

    Yosuke Higuchi, Brian Yan, Siddhant Arora, Tetsuji Ogawa, Tetsunori Kobayashi, and Shinji Watanabe. 2022. BERT Meets CTC: New Formulation of End-to-End Speech Recognition with Pre-trained Masked Language Model. In Findings of the Association for Computational Linguistics: EMNLP

  40. [42]

    Thomas Holtgraves. 2008. Automatic intention recognition in conversation processing. Journal of Memory and Language 58, 3 (2008), 627–645

  41. [43]

    Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processin...

  42. [44]

    Bo Hu, Kai Zhang, Yanghai Zhang, and Yuyang Ye. 2025. Adaptive Multimodal Fusion: Dynamic Attention Allocation for Intent Recognition. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 39. 17267–17275

  43. [45]

    Haigen Hu, Xiaoyuan Wang, Yan Zhang, Qi Chen, and Qiu Guan. 2024. A comprehensive survey on contrastive learning. Neurocomputing 610 (2024), 128645

  44. [46]

    Shijue Huang, Libo Qin, Bingbing Wang, Geng Tu, and Ruifeng Xu. 2024. Sdif-da: A shallow-to-deep interaction framework with data augmentation for multi-modal intent detection. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)...

  45. [47]

    Xinyue Huang and Adriana Kovashka. 2016. Inferring visual persuasion via body language, setting, and deep features. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops . 73–79

  46. [48]

    Matthew Huggins, Sharifa Alghowinem, Sooyeon Jeong, Pedro Colon-Hernandez, Cynthia Breazeal, and Hae Won Park. 2021. Practical guidelines for intent recognition: Bert with minimal training data evaluated in real-world hri application. In Proceedings of the 2021 ACM/IEEE Intern...

  47. [49]

    Oluwagbenga Paul Idowu, Ademola Enitan Ilesanmi, Xiangxin Li, Oluwarotimi Williams Samuel, Peng Fang, and Guanglin Li. 2021. An integrated deep learning model for motor intention recognition of multi-class EEG Signals in upper limb amputees. Computer Methods and Programs in Bi...

  48. [50]

    Siddarth Jain and Brenna Argall. 2018. Recursive bayesian human intent recognition in shared-control robotics. In 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 3905–3912

  49. [51]

    Siddarth Jain and Brenna Argall. 2019. Probabilistic human intent recognition for shared autonomy in assistive robotics. ACM Transactions on Human-Robot Interaction (THRI) 9, 1 (2019), 1–23

  50. [52]

    Dietmar Jannach and Markus Zanker. 2024. A survey on intent-aware recommender systems. ACM Transactions on Recommender Systems 3, 2 (2024), 1–32

  51. [54]

    Menglin Jia, Zuxuan Wu, Austin Reiter, Claire Cardie, Serge Belongie, and Ser-Nam Lim. 2021. Intentonomy: a dataset and study towards human intent understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 12986–12996

  52. [55]

    Guoqian Jiang, Kunyu Wang, Qun He, and Ping Xie. 2024. E2FNet: An EEG- and EMG-Based Fusion Network for Hand Motion Intention Recognition. IEEE Sensors Journal 24, 22 (2024), 38417–38428. doi:10.1109/JSEN.2024.3471894

  53. [56]

    Yidi Jiang, Bidisha Sharma, Maulik Madhavi, and Haizhou Li. 2021. Knowledge distillation from bert transformer to speech transformer for intent classification. In Proc. Interspeech 2021 (2021), 4713–4717

  54. [57]

    Di Jin, Shuyang Gao, Seokhwan Kim, Yang Liu, and Dilek Hakkani-Tür. 2022. Towards textual out-of-domain detection without in-domain labels. IEEE/ACM Transactions on Audio, Speech, and Language Processing 30 (2022), 1386–1395

  55. [58]

    Jungseock Joo, Weixin Li, Francis F Steen, and Song-Chun Zhu. 2014. Visual persuasion: Inferring communicative intents of images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 216–223

  56. [59]

    Jungseock Joo, Francis F Steen, and Song-Chun Zhu. 2015. Automated facial trait judgment and election outcome prediction: Social dimensions of face. In Proceedings of the IEEE International Conference on Computer Vision . 3712–3720

  57. [60]

    Gonuguntla, K.C

    Jun-Su Kang, Ukeob Park, V. Gonuguntla, K.C. Veluvolu, and Minho Lee. 2015. Human implicit intent recognition based on the phase synchrony of EEG signals. Pattern Recognition Letters 66 (2015), 144–152. doi:10.1016/j.patrec.2015.06.013

  58. [61]

    Yoon Kim. 2014. Convolutional Neural Networks for Sentence Classification. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 1746–1751. doi:10.3115/v1/D14-1181

  59. [62]

    Julia Kruk, Jonah Lubin, Karan Sikka, Xiao Lin, Dan Jurafsky, and Ajay Divakaran. 2019. Integrating Text and Image: Determining Multimodal Document Intent in Instagram Posts. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th ...

  60. [63]

    Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K

    Stefan Larson, Anish Mahendran, Joseph J. Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael A. Laurenzano, Lingjia Tang, and Jason Mars. 2019. An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction. In Proceed...

  61. [64]

    Gregory Lemasurier, Gal Bejerano, Victoria Albanese, Jenna Parrillo, Holly A Yanco, Nicholas Amerson, Rebecca Hetrick, and Elizabeth Phillips

  62. [65]

    Haoyang Li, Xin Wang, Ziwei Zhang, Jianxin Ma, Peng Cui, and Wenwu Zhu. 2021. Intention-aware sequential recommendation with structured intent transition. IEEE Transactions on Knowledge and Data Engineering 34, 11 (2021), 5403–5414. 30 J. Zhao et al

  63. [66]

    Jiapeng Li, Ping Wei, Wenjuan Han, and Lifeng Fan. 2023. Intentqa: Context-aware video intent reasoning. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 11963–11974

  64. [67]

    Jiahao Nick Li, Yan Xu, Tovi Grossman, Stephanie Santosa, and Michelle Li. 2024. Omniactions: Predicting digital actions in response to real-world multimodal sensory inputs with llms. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems . 1–22

  65. [68]

    Leyan Li, Rennong Yang, Maolong Lv, Ao Wu, and Zilong Zhao. 2024. From Behavior to Natural Language: Generative Approach for Unmanned Aerial Vehicle Intent Recognition. IEEE Transactions on Artificial Intelligence 5, 12 (2024), 6196–6209. doi:10.1109/TAI.2024.3376510

  66. [69]

    Mingrui Li, Zuoxu Wang, Fan Li, and Jihong Liu. 2025. A multi-task engineering design intention recognition approach based on Vision Transformer and EEG data. Advanced Engineering Informatics 65 (2025), 103353. doi:10.1016/j.aei.2025.103353

  67. [70]

    Tingyu Li, Junpeng Bao, Jiaqi Qin, Yuping Liang, Ruijiang Zhang, and Jason Wang. 2024. Multi-modal intent detection with lvamoe: the language-visual-audio mixture of experts. In 2024 IEEE International Conference on Multimedia and Expo (ICME) . IEEE, 1–6

  68. [71]

    Wenju Li, Yue Ma, Keyong Shao, Zhengkun Yi, Wujing Cao, Meng Yin, Tiantian Xu, and Xinyu Wu. 2024. The Human–Machine Interface Design Based on sEMG and Motor Imagery EEG for Lower Limb Exoskeleton Assistance System. IEEE Transactions on Instrumentation and Measurement 73 (2024...

  69. [72]

    Xinglin Li, Hanhui Deng, Jinhui Ouyang, Huayan Wan, Weiren Yu, and Di Wu. 2024. Act as What You Think: Towards Personalized EEG Interaction Through Attentional and Embedded LSTM Learning. IEEE Transactions on Mobile Computing 23, 5 (2024), 3741–3753. doi:10.1109/TMC.2023.3283022

  70. [73]

    Yanen Li, Bo-June Paul Hsu, and ChengXiang Zhai. 2013. Unsupervised identification of synonymous query intent templates for attribute intents. In Proceedings of the 22nd ACM International Conference on Information & Knowledge Management . 2029–2038

  71. [74]

    Yurong Li, Hao Yang, Jixiang Li, Dongyi Chen, and Min Du. 2020. EEG-based intention recognition with deep recurrent-convolution neural network: Performance and channel selection by Grad-CAM. Neurocomputing 415 (2020), 225–233. doi:10.1016/j.neucom.2020.07.072

  72. [75]

    Zhipeng Li, Binglin Wu, Yingyi Zhang, Xianneng Li, Kai Li, and Weizhi Chen. 2025. CuSMer: Multimodal Intent Recognition in Customer Service via Data Augment and LLM Merge. In Companion Proceedings of the ACM on Web Conference 2025 . 3058–3062

  73. [78]

    Jinggui Liang, Lizi Liao, Hao Fei, and Jing Jiang. 2024. Synergizing Large Language Models and Pre-Trained Smaller Models for Conversational Intent Discovery. In Findings of the Association for Computational Linguistics: ACL 2024 . 14133–14147. doi:10.18653/v1/2024.findings-acl.840

  74. [79]

    Bing Liu and Ian Lane. 2016. Attention-Based Recurrent Neural Network Models for Joint Intent Detection and Slot Filling. In Proc. Interspeech. 685–689

  75. [80]

    Junhua Liu, Tan Keat, Bin Fu, and Kwan Hui Lim. 2024. LARA: Linguistic-Adaptive Retrieval-Augmentation for Multi-Turn Intent Classification. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track . 1096–1106

  76. [83]

    Junhua Liu, Yong Keat Tan, Bin Fu, and Kwan Hui Lim. 2024. Intent-Aware Dialogue Generation and Multi-Task Contrastive Learning for Multi-Turn Intent Classification. arXiv preprint arXiv:2411.14252 (2024)

  77. [84]

    Rui Liu, Haolin Zuo, Zheng Lian, Xiaofen Xing, Björn W Schuller, and Haizhou Li. 2024. Emotion and intent joint understanding in multimodal conversation: A benchmarking dataset. arXiv preprint arXiv:2407.02751 (2024)

  78. [85]

    Yingxin Liu, Xinbin Liang, Yang Yu, Jianxiang Sun, Jiayao Hu, Yadong Liu, Ling-Li Zeng, Zongtan Zhou, and Dewen Hu. 2025. Recognizing drivers’ turning intentions with EEG and eye movement. Biomedical Signal Processing and Control 101 (2025), 107218. doi:10.1016/j.bspc.2024.107218

  79. [86]

    Zuojun Liu, Wei Lin, Yanli Geng, and Peng Yang. 2017. Intent Pattern Recognition of Lower-Limb Motion Based on Mechanical Sensors. IEEE/CAA Journal of Automatica Sinica 4, 4 (2017), 651–660. doi:10.1109/JAS.2017.7510619

  80. [87]

    Dylan P Losey, Craig G McDonald, Edoardo Battaglia, and Marcia K O’Malley. 2018. A review of intent detection, arbitration, and communication aspects of shared control for physical human–robot interaction. Applied Mechanics Reviews 70, 1 (2018), 010804

  81. [88]

    Adyasha Maharana, Quan Tran, Franck Dernoncourt, Seunghyun Yoon, Trung Bui, Walter Chang, and Mohit Bansal. 2022. Multimodal Intent Discovery from Livestream Videos. In Findings of the Association for Computational Linguistics: NAACL 2022 . 476–489

  82. [89]

    Andrew McCallum, Kamal Nigam, et al. 1998. A comparison of event models for naive bayes text classification. In AAAI-98 Workshop on Learning for Text Categorization, Vol. 752. 41–48

  83. [90]

    Grégoire Mesnil, Xiaodong He, Li Deng, and Yoshua Bengio. 2013. Investigation of recurrent-neural-network architectures and learning methods for spoken language understanding.. In Interspeech. 3771–3775

  84. [91]

    Trisha Mittal, Sanjoy Chowdhury, Pooja Guhan, Snikitha Chelluri, and Dinesh Manocha. 2024. Towards Determining Perceived Audience Intent for Multimodal Social Media Posts Using the Theory of Reasoned Action. Scientific Reports 14, 1 (2024), 10606. doi:10.1038/s41598-024-60299-w

  85. [92]

    Andrew Ng and Michael Jordan. 2001. On discriminative vs. generative classifiers: A comparison of logistic regression and naive bayes. Advances in Neural Information Processing Systems 14 (2001). Deep Learning Approaches for Multimodal Intent Recognition: A Survey 31

  86. [93]

    Phuong Ngoc-Duy Nguyen and Huan Hong Nguyen. 2024. Unveiling the link between digital entrepreneurship education and intention among university students in an emerging economy. Technological Forecasting and Social Change 203 (2024), 123330

  87. [94]

    Quynh-Mai Thi Nguyen, Lan-Nhi Thi Nguyen, and Cam-Van Thi Nguyen. 2024. TECO: Improving Multimodal Intent Recognition with Text Enhancement through Commonsense Knowledge Extraction. (2024), 533–541

  88. [95]

    Kawsar Noor, Katherine Smith, Jade O’Connell, Niamh Ingram, Baptiste Briot Ribyere, Tom Searle, Wai Keong Wong, and Richard J Dobson. 2024. Detecting Clinical Intent in Electronic Healthcare Records in a UK National Healthcare Hospital. In 2024 IEEE 12th International Conferen...

  89. [96]

    Innocent Otache, James Edomwonyi Edopkolor, Idris Ahmed Sani, and Kadiri Umar. 2024. Entrepreneurship education and entrepreneurial intentions: Do entrepreneurial self-efficacy, alertness and opportunity recognition matter? The International Journal of Management Education 22,...

  90. [97]

    Srinivas Bangalore Padmasundari and Srinivas Bangalore. 2018. Intent discovery through unsupervised semantic text clustering. InProc. Interspeech, Vol. 2018. 606–610

  91. [98]

    Gyutae Park, Ingeol Baek, ByeongJeong Kim, Joongbo Shin, and Hwanhee Lee. 2024. Dynamic Label Name Refinement for Few-Shot Dialogue Intent Classification. arXiv preprint arXiv:2412.15603 (2024)

  92. [99]

    Ukeob Park, Rammohan Mallipeddi, and Minho Lee. 2014. Human Implicit Intent Discrimination Using EEG and Eye Movement. In Neural Information Processing. Cham, 11–18

  93. [100]

    Max Pascher, Uwe Gruenefeld, Stefan Schneegass, and Jens Gerken. 2023. How to communicate robot motion intent: A scoping review. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–17

  94. [101]

    Kate Pearce, Sharifa Alghowinem, and Cynthia Breazeal. 2023. Build-a-bot: teaching conversational ai using a transformer-based intent recognition and question answering architecture. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 37. 16025–16032

  95. [102]

    Claudio Pinhanez, Paulo Cavalin, Victor Henrique Alves Ribeiro, Ana Appel, Heloisa Candello, Julio Nogima, Mauro Pichiliani, Melina Guerra, Maira de Bayser, Gabriel Malfatti, et al. 2021. Using meta-knowledge mined from identifiers to improve intent recognition in conversation...

  96. [103]

    Hamed Pirsiavash, Carl Vondrick, and Antonio Torralba. 2014. Inferring the why in images. arXiv preprint arXiv:1406.5472 2 (2014)

  97. [104]

    Libo Qin, Qiguang Chen, Jingxuan Zhou, Jin Wang, Hao Fei, Wanxiang Che, and Min Li. 2025. Divide-Solve-Combine: An Interpretable and Accurate Prompting Framework for Zero-shot Multi-Intent Detection. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 39. 2...

  98. [105]

    Quintero, I

    R. Quintero, I. Parra, J. Lorenzo, D. Fernández-Llorca, and M. A. Sotelo. 2017. Pedestrian Intention Recognition by Means of a Hidden Markov Model and Body Language. In 2017 IEEE 20th International Conference on Intelligent Transportation Systems (ITSC) . 1–7. doi:10.1109/ITSC...

  99. [106]

    Nugent, Jun Liu, and Liming Chen

    Joseph Rafferty, Chris D. Nugent, Jun Liu, and Liming Chen. 2017. From Activity Recognition to Intention Recognition for Assisted Living within Smart Homes. IEEE Transactions on Human-Machine Systems 47, 3 (2017), 368–379. doi:10.1109/THMS.2016.2641388

  100. [107]

    Swayambhu Nath Ray, Minhua Wu, Anirudh Raju, Pegah Ghahremani, Raghavendra Bilgi, Milind Rao, Harish Arsikere, Ariya Rastrow, Andreas Stolcke, and Jasha Droppo. 2021. Listen with intent: Improving speech recognition with audio-to-intent front-end. arXiv preprint arXiv:2105.070...

  101. [108]

    Juan A Rodriguez, Nicholas Botzer, David Vazquez, Christopher Pal, Marco Pedersoli, and Issam Laradji. 2024. IntentGPT: Few-shot Intent Discovery with Large Language Models. arXiv preprint arXiv:2411.10670 (2024)

  102. [109]

    Schultz and Saskia Kaiser

    Carsten D. Schultz and Saskia Kaiser. 2025. Consumer Value Dimensions in Conversational and Mobile Commerce. Journal of Marketing Analytics (2025), 1–19. doi:10.1057/s41270-025-00383-w

  103. [110]

    Jetze Schuurmans and Flavius Frasincar. 2019. Intent classification for dialogue utterances. IEEE Intelligent Systems 35, 1 (2019), 82–88

  104. [111]

    Mansi Sharma, Shuang Chen, Philipp Müller, Maurice Rekrut, and Antonio Krüger. 2023. Implicit Search Intent Recognition using EEG and Eye Tracking: Novel Dataset and Cross-User Prediction. In Proceedings of the 25th International Conference on Multimodal Interaction . 345–354

  105. [112]

    Neha Sharma, Chhavi Dhiman, and Sreedevi Indu. 2025. Predicting pedestrian intentions with multimodal IntentFormer: A Co-learning approach. Pattern Recognition 161 (2025), 111205

  106. [113]

    Yaomin Shen, Xiaojian Lin, and Wei Fan. 2025. A-MESS: Anchor based Multimodal Embedding with Semantic Synchronization for Multimodal Intent Recognition. arXiv preprint arXiv:2503.19474 (2025)

  107. [114]

    QingHongYa Shi, Mang Ye, Wenke Huang, Weijian Ruan, and Bo Du. 2024. Label-Aware Calibration and Relation-Preserving in Visual Intention Understanding. IEEE Transactions on Image Processing 33 (2024), 2627–2638

  108. [115]

    QingHongYa Shi, Mang Ye, Ziyi Zhang, and Bo Du. 2023. Learnable hierarchical label embedding and grouping for visual intention understanding. IEEE Transactions on Affective Computing 14, 4 (2023), 3218–3230

  109. [117]

    Jorge Sinval, Pedro Oliveira, Filipa Novais, Carla Maria Almeida, and Diogo Telles-Correia. 2024. Correlates of burnout and dropout intentions in medical students: A cross-sectional study. Journal of Affective Disorders 364 (2024), 221–230

  110. [118]

    Balazs, and Juan D

    Gino Slanzi, Jorge A. Balazs, and Juan D. Velásquez. 2017. Combining eye tracking, pupil dilation and EEG analysis for predicting web users click intention. Information Fusion 35 (2017), 51–57. doi:10.1016/j.inffus.2016.09.003 32 J. Zhao et al

  111. [119]

    Mohammad Soleymani, Michael Riegler, and Pål Halvorsen. 2017. Multimodal analysis of image search intent: Intent recognition in image search from user behavior and visual content. In Proceedings of the 2017 ACM on International Conference on Multimedia Retrieval . 251–259

  112. [120]

    Jinwang Song, Zhongtian Hua, Hongying Zan, Yingjie Han, and Min Peng. 2025. Optimizing Discriminative Vision-Language Models for Efficient Multimodal Intent Recognition. In Companion Proceedings of the ACM on Web Conference 2025 . 3063–3067

  113. [121]

    Kaili Sun, Zhiwen Xie, Mang Ye, and Huyin Zhang. 2024. Contextual augmented global contrast for multimodal intent recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 26963–26973

  114. [122]

    Stuart Synakowski, Qianli Feng, and Aleix Martinez. 2021. Adding knowledge to unsupervised algorithms for the recognition of intent.International journal of computer vision 129, 4 (2021), 942–959

  115. [123]

    Xianlun Tang, Tianzhu Wang, Xingchen Li, Wenbin Zhu, Xinyi Hong, and Xinbo Gao. 2024. Temporal Fusion Dynamically Separable Graph Convolutional Network for EEG Motion Intention Decoding Based on Source Information Extraction. IEEE Transactions on Instrumentation and Measuremen...

  116. [125]

    Yusheng Tian and Philip John Gorinski. 2020. Improving End-to-End Speech-to-Intent Classification with Reptile. InProc. Interspeech 2020. 891–895

  117. [126]

    Susanne Trick, Dorothea Koert, Jan Peters, and Constantin A Rothkopf. 2019. Multimodal uncertainty reduction for intention recognition in human-robot interaction. In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 7009–7016

  118. [127]

    Susanne Trick, Vilja Lott, Lisa Scherf, Constantin A Rothkopf, and Dorothea Koert. 2023. What can i help you with: Towards task-independent detection of intentions for interaction in a human-robot environment. In 2023 32nd IEEE International Conference on Robot and Human Inter...

  119. [128]

    Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9, 86 (2008), 2579–2605

  120. [129]

    Dimitrios Varytimidis, Fernando Alonso-Fernandez, Boris Duran, and Cristofer Englund. 2018. Action and Intention Recognition of Pedestrians in Urban Traffic. In 2018 14th International Conference on Signal-Image Technology & Internet-Based Systems (SITIS) . 676–682. doi:10.110...

  121. [130]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017)

  122. [131]

    Carl Vondrick, Deniz Oktay, Hamed Pirsiavash, and Antonio Torralba. 2016. Predicting motivations of actions by leveraging text. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 2997–3005

  123. [132]

    Binglu Wang, Kang Yang, Yongqiang Zhao, Teng Long, and Xuelong Li. 2023. Prototype-based intent perception. IEEE Transactions on Multimedia 25 (2023), 8308–8319

  124. [133]

    Jinpeng Wang, Gao Cong, Xin Zhao, and Xiaoming Li. 2015. Mining user intents in twitter: A semi-supervised approach to inferring intent categories for tweets. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 29

  125. [136]

    Shijie Wang, Zhiqiang Pu, Yi Pan, Boyin Liu, Hao Ma, and Jianqiang Yi. 2024. Long-Term and Short-Term Opponent Intention Inference for Football Multiplayer Policy Learning. IEEE Transactions on Cognitive and Developmental Systems 16, 6 (2024), 2055–2069

  126. [137]

    Yingzhi Wang, Abdelmoumene Boumadane, and Abdelwahab Heba. 2021. A fine-tuned wav2vec 2.0/hubert benchmark for speech emotion recognition, speaker verification and spoken language understanding. arXiv preprint arXiv:2111.02735 (2021)

  127. [138]

    Yaqing Wang, Song Wang, Yanyan Li, and Dejing Dou. 2022. Recognizing Medical Search Query Intent by Few-shot Learning. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval . 502–512. doi:10.1145/3477495.3531789

  128. [139]

    Christopher Yee Wong, Lucas Vergez, and Wael Suleiman. 2024. Vision- and Tactile-Based Continuous Multimodal Intention and Attention Recognition for Safer Physical Human–Robot Interaction. IEEE Transactions on Automation Science and Engineering 21, 3 (2024), 3205–3215

  129. [140]

    Jiaying Wu, Fanxiao Li, Min-Yen Kan, and Bryan Hooi. 2025. Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models. arXiv preprint arXiv:2505.15489 (2025)

  130. [141]

    Ting-Wei Wu, Ruolin Su, and Biing Juang. 2021. A label-aware BERT attention network for zero-shot multi-intent detection in spoken language understanding. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . 4884–4896

  131. [142]

    Dongfang Xu and Qining Wang. 2021. Noninvasive human-prosthesis interfaces for locomotion intent recognition: A review. Cyborg and Bionic Systems (2021)

  132. [144]

    Puyang Xu and Ruhi Sarikaya. 2013. Convolutional neural network based triangular crf for joint intent detection and slot filling. In 2013 Ieee Workshop on Automatic Speech Recognition and Understanding . IEEE, 78–83

  133. [145]

    Yijing Xu, Shifan Yu, Lei Liu, Wansheng Lin, Zhicheng Cao, Yu Hu, Jiming Duan, Zijian Huang, Chao Wei, Ziquan Guo, Tingzhu Wu, Zhong Chen, Qingliang Liao, Yuanjin Zheng, and Xinqin Liao. 2024. In-Sensor Touch Analysis for Intent Recognition. Advanced Functional Materials 34, 5...

  134. [146]

    Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association...

  135. [147]

    Weidong Yan, Jingyu Liu, Jie Luo, Wenkang Liu, Yulan Ma, and Qinge Zhang. 2025. A Temporal–Spatial Embedding and Dynamic Aggregation Network With Adaptive Weighting Spectrum for EEG Motion Intention Recognition. IEEE Transactions on Instrumentation and Measurement 74 (2025), 1–11

  136. [148]

    Bo Yang, Xinxing Chen, Xiling Xiao, Pei Yan, Yasuhisa Hasegawa, and Jian Huang. 2023. Gaze and Environmental Context-Guided Deep Neural Network and Sequential Decision Fusion for Grasp Intention Recognition. IEEE Transactions on Neural Systems and Rehabilitation Engineering 31...

  137. [149]

    Qu Yang, Qinghongya Shi, Tongxin Wang, and Mang Ye. 2025. Uncertain multimodal intention and emotion understanding in the wild. In Proceedings of the Computer Vision and Pattern Recognition Conference . 24700–24709

  138. [150]

    Qu Yang, Mang Ye, and Dacheng Tao. 2024. Synergy of sight and semantics: visual intention understanding with clip. In European Conference on Computer Vision. Springer, 144–160

  139. [151]

    Kaisheng Yao, Geoffrey Zweig, Mei-Yuh Hwang, Yangyang Shi, and Dong Yu. 2013. Recurrent neural networks for language understanding.. In Interspeech. 2524–2528

  140. [152]

    Mang Ye, Qinghongya Shi, Kehua Su, and Bo Du. 2023. Cross-modality pyramid alignment for visual intention understanding. IEEE Transactions on Image Processing 32 (2023), 2190–2201

  141. [153]

    Hao Yu, Jesujoba O Alabi, Andiswa Bukula, Jian Yun Zhuang, En-Shiun Annie Lee, Tadesse Kebede Guge, Israel Abebe Azime, Happy Buzaaba, Blessing Kudzaishe Sibanda, Godson K Kalipe, et al. 2025. INJONGO: A Multicultural Intent Detection and Slot-filling Dataset for 16 African La...

  142. [154]

    Habeeb Yusuf, Arthur Money, and Damon Daylamani-Zad. 2025. Pedagogical AI Conversational Agents in Higher Education: A Conceptual Framework and Survey of the State of the Art. Educational Technology Research and Development 73, 2 (2025), 815–874. doi:10.1007/s11423-025- 10447-4

  143. [155]

    Xiaoxue Zang, Abhinav Rastogi, Srinivas Sunkara, Raghav Gupta, Jianguo Zhang, and Jindong Chen. 2020. MultiWOZ 2.2 : A Dialogue Dataset with Additional Annotation Corrections and State Tracking Baselines. In Proceedings of the 2nd Workshop on Natural Language Processing for Co...

  144. [156]

    Li-Ming Zhan, Haowen Liang, Bo Liu, Lu Fan, Xiao-Ming Wu, and Albert YS Lam. 2021. Out-of-Scope Intent Detection with Self-Supervision and Discriminative Training. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Internati...

  145. [157]

    Bolin Zhang, Zhiying Tu, Shaoshi Hang, Dianhui Chu, and Xiaofei Xu. 2023. Conco-ernie: Complex user intent detect model for smart healthcare cognitive bot. ACM Transactions on Internet Technology 23, 1 (2023), 1–24

  146. [158]

    Zhang, K

    D. Zhang, K. Chen, D. Jian, L. Yao, S. Wang, and P. Li. 2019. Learning Attentional Temporal Cues of Brainwaves with Spatial Embedding for Motion Intent Detection. In 2019 IEEE International Conference on Data Mining (ICDM) . IEEE, 1450–1455. doi:10.1109/ICDM.2019.00189

  147. [159]

    Dalin Zhang, Lina Yao, Xiang Zhang, Sen Wang, Weitong Chen, Robert Boots, and Boualem Benatallah. 2018. Cascade and Parallel Convolutional Recurrent Neural Networks on EEG-based Intention Recognition for Brain Computer Interface. Proceedings of the AAAI Conference on Artificia...

  148. [160]

    Fan Zhang and He Huang. 2013. Source Selection for Real-Time User Intent Recognition toward Volitional Control of Artificial Legs. IEEE Journal of Biomedical and Health Informatics 17, 5 (2013), 907–914. doi:10.1109/JBHI.2012.2236563

  149. [161]

    Hanlei Zhang, Xiaoteng Li, Hua Xu, Panpan Zhang, Kang Zhao, and Kai Gao. 2021. TEXTOIR: An Integrated and Visualized Platform for Text Open Intent Recognition. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International...

  150. [162]

    Hanlei Zhang, Xin Wang, Hua Xu, Qianrui Zhou, Kai Gao, Jianhua Su, jinyue Zhao, Wenrui Li, and Yanting Chen. 2024. MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations. In The Twelfth International Conference...

  151. [163]

    Hanlei Zhang, Hua Xu, Xin Wang, Qianrui Zhou, Shaojie Zhao, and Jiayan Teng. 2022. Mintrec: A new dataset for multimodal intent recognition. In Proceedings of the 30th ACM international conference on multimedia . 1688–1697

  152. [165]

    Lu Zhang, Jialie Shen, Jian Zhang, Jingsong Xu, Zhibin Li, Yazhou Yao, and Litao Yu. 2021. Multimodal marketing intent analysis for effective targeted advertising. IEEE Transactions on Multimedia 24 (2021), 1830–1843

  153. [166]

    Lu Zhang, Jialie Shen, Jian Zhang, Jingsong Xu, Zhibin Li, Yazhou Yao, and Litao Yu. 2022. Multimodal Marketing Intent Analysis for Effective Targeted Advertising. IEEE Transactions on Multimedia 24 (2022), 1830–1843. doi:10.1109/TMM.2021.3073267

  154. [167]

    Lu Zhang, Jian Zhang, Zhibin Li, and Jingsong Xu. 2020. Towards better graph representation: Two-branch collaborative graph neural networks for multimodal marketing intention detection. In 2020 IEEE International Conference on Multimedia and Expo (ICME) . IEEE, 1–6

  155. [168]

    Rongzheng Zhang, Wanghongjie Qiu, Jianuo Qiu, Yuqin Guo, Chengxiao Dong, Tuo Zhang, Juan Yi, Chaoyang Song, Harry Asada, and Fang Wan. 2025. MultiModal Intention Recognition Combining Head Motion and Throat Vibration for Underwater Superlimbs. IEEE Transactions on 34 J. Zhao e...

  156. [169]

    Shun Zhang, Yan Chaoran, Jian Yang, Jiaheng Liu, Ying Mo, Jiaqi Bai, Tongliang Li, and Zhoujun Li. 2024. Towards Real-world Scenario: Imbalanced New Intent Discovery. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pape...

  157. [170]

    Wenxin Zhang, Hao Qi, Shutong Wang, Ziqi Lin, and Binglu Wang. 2024. Content and Relation Fuzzy Mitigation Framework for Intent Perception. IEEE Transactions on Fuzzy Systems (2024), 1–14. doi:10.1109/TFUZZ.2024.3476948

  158. [171]

    Xin Zhang, Fei Cai, Xuejun Hu, Jianming Zheng, and Honghui Chen. 2022. A contrastive learning-based task adaptation model for few-shot intent recognition. Information Processing & Management 59, 3 (2022), 102863

  159. [172]

    Xiyuan Zhao, Huijun Li, Tianyuan Miao, Xianyi Zhu, Zhikai Wei, Lifen Tan, and Aiguo Song. 2024. Learning Multimodal Confidence for Intention Recognition in Human-Robot Interaction. IEEE Robotics and Automation Letters 9, 9 (2024), 7819–7826. doi:10.1109/LRA.2024.3432352

  160. [173]

    Yinhe Zheng, Guanyi Chen, and Minlie Huang. 2020. Out-of-domain detection for natural language understanding in dialog systems. IEEE/ACM Transactions on Audio, Speech, and Language Processing 28 (2020), 1198–1209

  161. [174]

    Qianrui Zhou, Hua Xu, Hao Li, Hanlei Zhang, Xiaohan Zhang, Yifan Wang, and Kai Gao. 2024. Token-level contrastive learning with modality-aware prompting for multimodal intent recognition. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 17114–17122

  162. [175]

    Yunhua Zhou, Peiju Liu, and Xipeng Qiu. 2022. KNN-contrastive learning for out-of-domain intent classification. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 5129–5141

  163. [176]

    Yuxi Zhou, Xiujie Wang, Jianhua Zhang, Jiajia Wang, Jie Yu, Hao Zhou, Yi Gao, and Shengyong Chen. 2024. Intentional evolutionary learning for untrimmed videos with long tail distribution. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 7713–7721

  164. [177]

    Zhihong Zhu, Xuxin Cheng, Zhaorun Chen, Yuyan Chen, Yunyan Zhang, Xian Wu, Yefeng Zheng, and Bowen Xing. 2024. InMu-Net: advancing multi-modal intent detection via information bottleneck and multi-sensory processing. In Proceedings of the 32nd ACM International Conference on M...

  165. [2020]

    In Proceedings of the Twelfth Language Resources and Evaluation Conference

    MultiWOZ 2.1: A Consolidated Multi-Domain Dialogue Dataset with State Corrections and State Tracking Baselines. In Proceedings of the Twelfth Language Resources and Evaluation Conference . 422–428. https://aclanthology.org/2020.lrec-1.53/

  166. [2021]

    ACM Transactions on Human-Robot Interaction (THRI) 10, 4 (2021), 1–27

    Methods for expressing robot intent for human–robot collaboration in shared workspaces. ACM Transactions on Human-Robot Interaction (THRI) 10, 4 (2021), 1–27

  167. [2022]

    doi:10.18653/v1/2022.findings-emnlp.402 Deep Learning Approaches for Multimodal Intent Recognition: A Survey 29

    5486–5503. doi:10.18653/v1/2022.findings-emnlp.402 Deep Learning Approaches for Multimodal Intent Recognition: A Survey 29

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.