REVIEW 4 major objections 5 minor 3 cited by
Learning Sparsity for Effective and Efficient Music Performance Question Answering
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Sparsify, a sparse-learning pipeline for music performance question answering, reports state-of-the-art accuracy on two MUSIC-AVQA benchmarks while cutting training time by 28.32% and retaining 70–80% of full-data accuracy from a 25% key…
desk verdict Competent integration of three borrowed sparsification tricks with internally consistent numbers, but the SOTA accuracy claim is untestable without a dense Amuse baseline. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is three cooperating sparsification modules wrapped around a universal audio-visual-language encoder. Sparse Masking randomly drops 50% of video patches and mel-spectrogram audio tokens during the first three epochs. Adaptive Sparse Merging ranks tokens by cross-modal attention scores, keeps the top-quartile tokens (selected by interquartile range), and merges the rest into the nearest key token by key-vector similarity. Sparse Subset Selection computes per-sample losses, splits samples into hard and easy groups, aggregates hard-sample scores across epochs with a decay ratio, and uses InfoBatch to rescale gradients and prune easy samples. Together they reduce token count and dataset size while preserving question-relevant information.
What would settle it
Run the same universal encoder with all three sparsification components disabled on both MUSIC-AVQA test sets and compare accuracy with and without sparsification; if the dense encoder reaches 81.75% or higher without masking, merging, or subset selection, then the reported state-of-the-art result is not caused by sparse learning.
Extended reading notes
Core claim
Sparsify is claimed to be the first Music AVQA method that explicitly builds sparsification into every stage of training. Its universal encoder—built on a Swin-V2 video backbone, an HTS-AT audio backbone, and a text transformer—is trained with random 50% masking of audio and visual patches, adaptive merging of low-salience tokens into nearby key tokens selected by cross-modal attention, and InfoBatch-guided pruning that keeps hard examples while rescaling gradients. The paper reports that this pipeline outperforms AVST, LAVisH, and DG-SCT on both MUSIC-AVQA (81.75 overall) and MUSIC-AVQA v2.0 (81.30 overall), with the largest margins on audio-visual question types. Training time for the v2.0 benchmark drops 28.32% compared to a dense variant with all three strategies disabled, and a selected 25% key subset still yields 60.17% accuracy for Sparsify and 55.21% for DG-SCT, about 74% of their full-data scores.
Load-bearing premise
The accuracy results assume the three sparsification strategies, not the underlying encoder, are what lifts Sparsify above the baselines; the paper never reports the dense encoder's test accuracy, so this attribution is unverified.
Editorial extensions
If this is right
- If the reported numbers hold, a Music AVQA model can beat strong published baselines while training on 25% of the data at 72% of the wall-clock time, so dense audio-visual inputs are not a prerequisite for accuracy.
- The key-subset selection transfers across at least two different model architectures, suggesting the selected samples carry dataset-level difficulty information rather than model-specific quirks.
- Sparse Masking plus Adaptive Sparse Merging improves audio-visual QA by large margins (up to +11.24% over DG-SCT on audio-visual questions), indicating that aggressive token reduction is compatible with fine-grained reasoning.
- Training time savings are additive: 50% masking in early epochs, token merging throughout, and gradient-rescaled sample pruning each contribute to the 28.32% reduction.
Reading between the lines
- The authors leave untested whether the accuracy gains come from sparsification or from the stronger universal encoder; an ablation that runs the dense Amuse encoder on the same benchmarks would isolate the cause.
- Because the key subset retains roughly the same fraction of performance for two different models, a promising extension is to use the algorithm to build a reusable coreset for other dense multimodal benchmarks, reusing the selected indices across architectures.
- The fixed 50% masking rate and the 0.618 decay ratio are tuned on Music AVQA; a testable extension would sweep these hyperparameters on datasets with different redundancy levels to see whether the efficiency gains hold.
- If sparse training truly improves accuracy by suppressing background clutter, the same masking and merging recipe could transfer to other dense continuous audio-visual tasks such as instrument counting in orchestras or action recognition in concerts.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Sparsify, a sparse learning framework for Music Performance Audio-Visual Question Answering (Music AVQA). Sparsify combines three sparsification strategies—random masking of audio/visual tokens, adaptive token merging based on cross-modal attention, and InfoBatch-style curriculum subset selection—integrated into the authors' Amuse universal encoder. The manuscript reports state-of-the-art accuracy on MUSIC-AVQA (81.75 overall) and MUSIC-AVQA v2.0 (81.30), a 28.32% training-time reduction versus a dense variant (124 vs. 173 hours), and a key-subset selection method that uses ~25% of the training data while retaining 70–80% of full-data accuracy across models.
Significance. If the claims are substantiated, Sparsify would offer a practical efficiency improvement for a challenging multimodal QA task, and the key-subset selection idea has potential value for data-efficient audio-visual learning. The paper's reported arithmetic is internally consistent (e.g., 60.17/0.7401 matches the 81.30 full-data accuracy), which lends some credibility. However, the experimental design currently does not isolate the contribution of sparsification from the choice of encoder, because the dense Amuse baseline is never evaluated for accuracy and no ablations are provided for the three proposed strategies. The key-subset experiment also lacks a random-subset control, leaving the selection algorithm's benefit unproven. These are load-bearing gaps in the current submission.
major comments (4)
- [§3.2, Table 1; §2.1] The state-of-the-art claim is not attributable to sparse learning because the paper never reports the accuracy of the dense Amuse variant (all three sparsification strategies disabled). The Universal Encoder in §2.1 is the authors' own Amuse model with Swin-V2 and HTS-AT backbones, while the baselines AVST, LAVisH, and DG-SCT use different backbones. Figure 5 reports only training time (124 vs. 173 hours), not accuracy. Without the dense-Amuse accuracy on both MUSIC-AVQA and MUSIC-AVQA v2.0, the gains in Table 1 could stem entirely from the encoder choice, making the abstract's 'maintaining accuracy' and the SOTA headline untestable. Please add the dense variant's accuracy on both datasets.
- [§2.2–2.4] No ablation isolates the contribution of Sparse Masking, Adaptive Sparse Merging, and InfoBatch. The three strategies are introduced as independent components, but all experiments combine them. To support the claim that each strategy contributes to accuracy or efficiency, report results for each strategy enabled individually and in pairs, in addition to the full combination and the dense baseline.
- [§3.3, Figure 4] The key-subset experiment does not demonstrate that the selection algorithm outperforms random subsetting. Comparing subset training to full-data training only shows that 25% of the data retains 74% of performance; a random 25% subset control is needed to show that the choice of samples matters. Moreover, the subset is trained for 1 warm-up + 15 epochs (Section 3.1) while the full-data training budget is not specified to be matched, so the reported retention ratio may reflect unequal compute rather than the selection algorithm's benefit.
- [Algorithm 1] Algorithm 1 is internally inconsistent. The variable t is described as a temporary count vector but is never reset between epochs, so EpochsList accumulates cumulative counts rather than per-epoch scores; the merge step then adds these cumulative vectors with weights w_g, which does not correspond to the text's description of 'scores aggregated by epoch.' As written, the algorithm is not reproducible. Please correct the pseudocode to match the implemented procedure and specify how scores are normalized across epochs.
minor comments (5)
- [Figure 1] The caption text 'QA with Dense AudiofromMUSIC-AVQA v2.0' has a missing space; 'Audiofrom' should be 'Audio from'.
- [§3.1 / §3.3] The text uses '10,819 samples' in Figure 4 and 'Num = 10,819 (i.e., the number of QA pairs)' in Section 3.1; please use consistent terminology, because QA pairs and samples may not be identical.
- [§2.3] The attention formula a = softmax(Q·K^T / sqrt(d)) V appears to define the full attention output, while the text says it evaluates token importance; clarify that a denotes the attention weights before the final multiplication by V.
- [Figure 4] The bar chart lacks axis labels; add both axis labels and numeric accuracy values to make the comparison readable.
- [Algorithm 1] The symbol N is used for the number of samples in Algorithm 1, while the later text uses Num for the key-subset size; rename one of these to avoid confusion.
Circularity Check
No circularity: the SOTA and efficiency claims are empirical measurements on external benchmarks; the missing dense-Amuse ablation is an experimental control gap, not a circular reduction.
full rationale
The paper's central claims are empirical results measured on the fixed external benchmarks MUSIC-AVQA and MUSIC-AVQA v2.0, not quantities derived from their own definitions. Table 1 directly reports accuracies for Sparsify and three published baselines; Figure 5 reports training time for Sparsify against a dense variant that disables all three sparsification strategies, providing a controlled comparison for the 28.32% efficiency gain. The key-subset experiment (Figure 4) reports measured accuracies on the selected subset versus full-data training. Although the Universal Encoder is the authors' own Amuse framework (cited as Diao et al., 2024), the paper does not derive its results from that citation; it runs experiments. The absence of a reported dense-Amuse accuracy means that the attribution of accuracy gains specifically to sparsification is not isolated from the encoder choice, but this is a missing ablation, not a case where a prediction is equivalent to an input by construction or where a load-bearing argument reduces to a self-citation. No equation equates a fitted or selected quantity with the reported outcome, and no claimed result is defined in terms of the target conclusion. The paper is self-contained against external benchmarks, so the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (8)
- Sparse masking ratio =
50%
- Masking schedule =
first 3 epochs only
- InfoBatch ratio =
0.5
- InfoBatch delta =
0.875
- Key-subset decay ratio r =
0.618
- Merge group size k =
3
- Key-subset size Num =
10,819 QA pairs (~25% of training data)
- Key-subset training budget =
1 warm-up + 15 epochs
assumptions (5)
- domain assumption Cross-modal attention scores a = softmax(QK^T/sqrt(d))V identify task-relevant tokens
- domain assumption InfoBatch loss-based difficulty scoring preserves training distribution statistics after pruning
- domain assumption Published baseline accuracies (AVST, LAVisH, DG-SCT) are directly comparable to Sparsify's numbers
- domain assumption Pretrained encoder backbones (Swin-V2, HTS-AT, standard question transformer) transfer to Music AVQA
- ad hoc to paper The three sparsification methods remain effective when combined with unchanged default hyperparameters
Cite this review
Pith. "Pith review of Learning Sparsity for Effective and Efficient Music Performance Question Answering." pith.science (2026). https://pith.science/paper/QHZLGHUC
@misc{pith2026250601319,
author = {Pith},
title = {Pith review of: Learning Sparsity for Effective and Efficient Music Performance Question Answering},
year = {2026},
howpublished = {\url{https://pith.science/paper/QHZLGHUC}},
note = {Machine review of arXiv:2506.01319}
}
read the original abstract
Music performances, characterized by dense and continuous audio as well as seamless audio-visual integration, present unique challenges for multimodal scene understanding and reasoning. Recent Music Performance Audio-Visual Question Answering (Music AVQA) datasets have been proposed to reflect these challenges, highlighting the continued need for more effective integration of audio-visual representations in complex question answering. However, existing Music AVQA methods often rely on dense and unoptimized representations, leading to inefficiencies in the isolation of key information, the reduction of redundancy, and the prioritization of critical samples. To address these challenges, we introduce Sparsify, a sparse learning framework specifically designed for Music AVQA. It integrates three sparsification strategies into an end-to-end pipeline and achieves state-of-the-art performance on the Music AVQA datasets. In addition, it reduces training time by 28.32% compared to its fully trained dense counterpart while maintaining accuracy, demonstrating clear efficiency gains. To further improve data efficiency, we propose a key-subset selection algorithm that selects and uses approximately 25% of MUSIC-AVQA v2.0 training data and retains 70-80% of full-data performance across models.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 3 Pith papers
-
FakeSV-VLM: Taming VLM for Detecting Fake Short-Video News via Progressive Mixture-Of-Experts Adapter
FakeSV-VLM reaches 90.22% and 89.30% accuracy on FakeSV and FakeTT by adding a two-stage MoE adapter and contrastive alignment to InternVL2.5-8B.
-
A Multimodal Deep Learning Framework for Early Diagnosis of Liver Cancer via Optimized BiLSTM-AM-VMD Architecture
The paper claims a BiLSTM-AM-VMD model achieves AUC 0.963 for early HCC diagnosis, but the evidence is undermined by contradictory dataset descriptions and missing artifacts.
-
Multi-Modal Machine Learning Framework for Predicting Early Recurrence of Brain Tumors Using MRI and Clinical Biomarkers
XGBoost combining MRI radiomics and clinical biomarkers reportedly reaches C-index 0.782 for early brain tumor recurrence, but the paper's methods describe a liver-cancer cohort and no evaluation of its claimed tempor...
Reference graph
Works this paper leans on
-
[1]
Asma Ben Abacha and Pierre Zweigenbaum. 2015. Means: A medical question-answering system combining nlp techniques and semantic web technologies. Information Processing & Management
2015
-
[2]
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In International Conference on Computer Vision
2015
-
[3]
Sagar S Arya, Sofia B Dias, Herbert F Jelinek, Leontios J Hadjileontiadis, and Anna-Maria Pappa. 2023. The convergence of traditional and digital biomarkers through ai-assisted biosensing: A new era in translational diagnostics? Biosensors and Bioelectronics
work page 2023
-
[4]
Alexandre Blanco-Gonzalez, Alfonso Cabezon, Alejandro Seco-Gonzalez, Daniel Conde-Torres, Paula Antelo-Riveiro, Angel Pineiro, and Rebeca Garcia-Fandino. 2023. The role of ai in drug discovery: challenges, opportunities, and strategies. Pharmaceuticals
work page 2023
-
[5]
Florian Bordes, Richard Yuanzhe Pang, Anurag Ajay, Alexander C Li, Adrien Bardes, Suzanne Petryk, Oscar Ma \ n as, Zhiqiu Lin, Anas Mahmoud, Bargav Jayaraman, et al. 2024. An introduction to vision-language modeling. arXiv preprint arXiv:2405.17247
arXiv 2024
-
[6]
Zal \'a n Borsos, Rapha \"e l Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al. 2023. Audiolm: a language modeling approach to audio generation. Transactions on Audio, Speech, and Language Processing
work page 2023
-
[7]
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman. 2020. Vggsound: A large-scale audio-visual dataset. In International Conference on Acoustics, Speech and Signal Processing
work page 2020
-
[8]
Ke Chen, Xingjian Du, Bilei Zhu, Zejun Ma, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2022 a . Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection. In International Conference on Acoustics, Speech and Signal Processing
work page 2022
Show all 82 references
-
[9]
Zhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma, Sameena Shah, and William Yang Wang. 2022 b . Convfinqa: Exploring the chain of numerical reasoning in conversational finance question answering. arXiv preprint arXiv:2210.03849
2022 arXiv
-
[10]
Suresh Dara, Swetha Dhamercherla, Surender Singh Jadav, CH Madhu Babu, and Mohamed Jawed Ahsan. 2022. Machine learning in drug discovery: a review. Artificial Intelligence Review
2022
-
[11]
Soham Deshmukh, Benjamin Elizalde, Rita Singh, and Huaming Wang. 2023. Pengi: An audio language model for audio tasks. Advances in Neural Information Processing Systems
2023
-
[12]
Xingjian Diao, Ming Cheng, and Shitong Cheng. 2023. Av-maskenhancer: Enhancing video representations through audio-visual masked autoencoder. In International Conference on Tools with Artificial Intelligence
2023
-
[13]
Xingjian Diao, Chunhui Zhang, Tingxuan Wu, Ming Cheng, Zhongyu Ouyang, Weiyi Wu, and Jiang Gui. 2024. Learning musical representations for music performance question answering. In Findings of the Association for Computational Linguistics: EMNLP
2024
-
[14]
Xingjian Diao, Chunhui Zhang, Weiyi Wu, Zhongyu Ouyang, Peijun Qing, Ming Cheng, Soroush Vosoughi, and Jiang Gui. 2025. Temporal working memory: Query-guided segment refinement for enhanced multimodal understanding. arXiv preprint arXiv:2502.06020
2025 arXiv
-
[15]
Haoyi Duan, Yan Xia, Zhou Mingze, Li Tang, Jieming Zhu, and Zhou Zhao. 2023. Cross-modal prompts: Adapting large pre-trained models for audio-visual downstream tasks. In Advances in Neural Information Processing Systems
2023
-
[16]
Fayek and Justin Johnson
Haytham M. Fayek and Justin Johnson. 2020. Temporal reasoning via audio question answering. Transactions on Audio, Speech, and Language Processing
2020
-
[17]
Chongyang Gao, Yiren Jian, Natalia Denisenko, Soroush Vosoughi, and VS Subrahmanian. 2024. Gem: generating engaging multimodal content. In International Joint Conference on Artificial Intelligence
2024
-
[18]
Kaixiong Gong, Kaituo Feng, Bohao Li, Yibing Wang, Mofan Cheng, Shijia Yang, Jiaming Han, Benyou Wang, Yutong Bai, Zhuoran Yang, et al. 2024. Av-odyssey bench: Can your multimodal llms really understand audio-visual information? arXiv preprint arXiv:2412.02611
2024 arXiv
-
[19]
Travis R Goodwin and Sanda M Harabagiu. 2016. Medical question answering for clinical decision support. In International on Conference on Information and Knowledge Management
2016
-
[20]
Yangfan He, Sida Li, Jianhui Wang, Kun Li, Xinyuan Song, Xinhang Yuan, Keqin Li, Kuan Lu, Menghao Huo, Jiaqi Chen, et al. 2025 a . Enhancing low-cost video editing with lightweight adaptors and temporal-aware inversion. arXiv preprint arXiv:2501.04606
2025 arXiv
-
[21]
Yangfan He, Jianhui Wang, Kun Li, Yijin Wang, Li Sun, Jun Yin, Miao Zhang, and Xueqian Wang. 2025 b . Enhancing intent understanding for ambiguous prompts through human-machine co-adaptation. arXiv preprint arXiv:2501.15167
2025
-
[22]
Panwen Hu, Nan Xiao, Feifei Li, Yongquan Chen, and Rui Huang. 2023. A reinforcement learning-based automatic video editing method using pre-trained vision-language model. In International Conference on Multimedia
2023
-
[23]
Allen H Huang, Hui Wang, and Yi Yang. 2023. Finbert: A large language model for extracting information from financial text. Contemporary Accounting Research
2023
-
[24]
Vladimir Iashin, Weidi Xie, Esa Rahtu, and Andrew Zisserman. 2022. Sparse in space and time: Audio-visual synchronisation with trainable selectors. arXiv preprint arXiv:2210.07055
2022 arXiv
-
[25]
Vladimir Iashin, Weidi Xie, Esa Rahtu, and Andrew Zisserman. 2024. Synchformer: Efficient synchronization from sparse cues. In International Conference on Acoustics, Speech and Signal Processing
2024
-
[26]
Yiren Jian, Chongyang Gao, and Soroush Vosoughi. 2023. Bootstrapping vision-language learning with decoupled language pre-training. In Advances in Neural Information Processing Systems
2023
-
[27]
Yiren Jian, Tingkai Liu, Yunzhe Tao, Chunhui Zhang, Soroush Vosoughi, and Hongxia Yang. 2024. Expedited training of visual conditioned language generation via redundancy reduction. In Annual Meeting of the Association for Computational Linguistics
2024
-
[28]
Balaram Yadav Kasula. 2023. Harnessing machine learning for personalized patient care. Transactions on Latest Trends in Artificial Intelligence
2023
-
[29]
Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar, et al. 2023. Voicebox: Text-guided multilingual universal speech generation at scale. Advances in Neural Information Processing Systems
2023
-
[30]
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg. 2018. Tvqa: Localized, compositional video question answering. In Conference on Empirical Methods in Natural Language Processing
2018
-
[31]
Bin Li and Hanjun Deng. 2023. Bilateral personalized dialogue generation with contrastive learning. Soft Computing
2023
-
[32]
Bin Li, Bin Sun, Shutao Li, Encheng Chen, Hongru Liu, Yixuan Weng, Yongping Bai, and Meiling Hu. 2024 a . Distinct but correct: generating diversified and entity-revised medical response. Science China Information Sciences
2024
-
[33]
Bowen Li, Xiaojuan Qi, Thomas Lukasiewicz, and Philip Torr. 2019. Controllable text-to-image generation. Advances in Neural Information Processing Systems
2019
-
[34]
Guangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu, Ji-Rong Wen, and Di Hu. 2022. Learning to answer questions in dynamic audio-visual scenarios. In Conference on Computer Vision and Pattern Recognition
2022
-
[35]
Shutao Li, Bin Li, Bin Sun, and Yixuan Weng. 2024 b . Towards visual-prompt temporal answer grounding in instructional video. Transactions on Pattern Analysis and Machine Intelligence
2024
-
[36]
Yanghao Li, Haoqi Fan, Ronghang Hu, Christoph Feichtenhofer, and Kaiming He. 2023 a . Scaling language-image pre-training via masking. In Conference on Computer Vision and Pattern Recognition
2023
-
[37]
Yinheng Li, Shaofei Wang, Han Ding, and Hang Chen. 2023 b . Large language models in finance: A survey. In International Conference on AI in Finance
2023
-
[38]
Jinhua Liang, Huan Zhang, Haohe Liu, Yin Cao, Qiuqiang Kong, Xubo Liu, Wenwu Wang, Mark D Plumbley, Huy Phan, and Emmanouil Benetos. 2024. Wavcraft: Audio editing and generation with large language models. arXiv preprint arXiv:2403.09527
2024 arXiv
-
[39]
Yan-Bo Lin, Yi-Lin Sung, Jie Lei, Mohit Bansal, and Gedas Bertasius. 2023. Vision transformers are parameter-efficient audio-visual learners. In Conference on Computer Vision and Pattern Recognition
2023
-
[40]
Samuel Lipping, Parthasaarathy Sudarsanam, Konstantinos Drossos, and Tuomas Virtanen. 2022. Clotho-aqa: A crowdsourced dataset for audio question answering. In European Signal Processing Conference
2022
-
[41]
Xiulong Liu, Zhikang Dong, and Peng Zhang. 2024. Tackling data bias in music-avqa: Crafting a balanced dataset for unbiased question-answering. In Winter Conference on Applications of Computer Vision
2024
-
[42]
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. 2025. Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision
2025
-
[43]
Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. 2022. Swin transformer v2: Scaling up capacity and resolution. In Conference on Computer Vision and Pattern Recognition
2022
-
[44]
Dongchan Min, Dong Bok Lee, Eunho Yang, and Sung Ju Hwang. 2021. Meta-stylespeech: Multi-speaker adaptive text-to-speech generation. In International Conference on Machine Learning
2021
-
[45]
Gianluca Monaci, Friedrich T Sommer, and Pierre Vandergheynst. 2008. Learning sparse generative models of audiovisual signals. In European Signal Processing Conference
2008
-
[46]
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng. 2011. Multimodal deep learning. In International Conference on Machine Learning
2011
-
[47]
Yingwei Pan, Yehao Li, Jianjie Luo, Jun Xu, Ting Yao, and Tao Mei. 2022. Auto-captions on gif: A large-scale video-sentence dataset for vision-language pre-training. In International Conference on Multimedia
2022
-
[48]
Qi Qian, Yuanhong Xu, and Juhua Hu. 2023. Intra-modal proxy learning for zero-shot visual categorization with clip. Advances in Neural Information Processing Systems
2023
-
[49]
Ziheng Qin, Kai Wang, Zangwei Zheng, Jianyang Gu, Xiangyu Peng, Zhaopan Xu, Daquan Zhou, Lei Shang, Baigui Sun, Xuansong Xie, and Yang You. 2024. Infobatch: Lossless training speed up by unbiased dynamic data pruning. In International Conference on Learning Representations
2024
-
[50]
Khyati Saini and Pardeep Singh. 2023. Evolution of financial question answering themes, challenges, and advances. In International Conference on Recent Innovations in Computing
2023
-
[51]
Yuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee, and Yan Yan. 2024. Llava-prumerge: Adaptive token reduction for efficient large multimodal models. arXiv preprint arXiv:2403.15388
2024
-
[52]
Peng Shen, Satoshi Tamura, and Satoru Hayamizu. 2013. Audio-visual interaction in sparse representation features for noise robust audio-visual speech recognition. In Auditory-Visual Speech Processing
2013
-
[53]
Fangxun Shu, Lei Zhang, Hao Jiang, and Cihang Xie. 2023. Audio-visual llm for video understanding. arXiv preprint arXiv:2312.06720
2023 arXiv
-
[54]
Teotino Gomes Soares, Azhari Azhari, Nur Rokhman, and E Wonarko. 2021. Education question answering systems: a survey. In International MultiConference of Engineers and Computer Scientists
2021
-
[55]
Salakhutdinov
Nitish Srivastava and Russ R. Salakhutdinov. 2012. Multimodal learning with deep boltzmann machines. In Advances in Neural Information Processing Systems
2012
-
[56]
Tim Steuer, Anna Filighera, and Thomas Tregel. 2022. Investigating educational and noneducational answer selection for educational question generation. IEEE Access
2022
-
[57]
Yi Su, Jisheng Bai, Qisheng Xu, Kele Xu, and Yong Dou. 2025. Audio-language models for audio-centric tasks: A survey. arXiv preprint arXiv:2501.15177
2025 arXiv
-
[58]
Deeksha Varshney, Aizan Zafar, Niranshu Kumar Behera, and Asif Ekbal. 2023. Knowledge graph assisted end-to-end medical dialog generation. Artificial Intelligence in Medicine
2023
-
[59]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems
2017
-
[60]
Shiru Wang, Yao Chen, Lesley A Jarvis, Yucheng Tang, David J Gladstone, Kimberley S Samkoe, Brian W Pogue, Petr Bruza, and Rongxiao Zhang. 2024. Robust real-time segmentation of bio-morphological features in human cherenkov imaging during radiotherapy via deep learning. arXiv ...
2024 arXiv
-
[61]
Yanbo J Wang, Yuming Li, Hui Qin, Yuhang Guan, and Sheng Chen. 2022. A novel deberta-based model for financial question answering task. arXiv preprint arXiv:2207.05875
2022 arXiv
-
[62]
Yuancheng Wang, Zeqian Ju, Xu Tan, Lei He, Zhizheng Wu, Jiang Bian, et al. 2023. Audit: Audio editing by following instructions with latent diffusion models. Advances in Neural Information Processing Systems
2023
-
[63]
Yake Wei, Di Hu, Yapeng Tian, and Xuelong Li. 2022. Learning in audio-visual context: A review, analysis, and new perspective. arXiv preprint arXiv:2208.09579
2022 arXiv
-
[64]
Yuxiang Wei, Anees Abrol, and Vince D Calhoun. 2025. Hierarchical spatio-temporal state-space modeling for fmri analysis. In International Conference on Research in Computational Molecular Biology
2025
-
[65]
Yuxiang Wei, Yuqian Chen, Tengfei Xue, Leo Zekelman, Nikos Makris, Yogesh Rathi, Weidong Cai, Fan Zhang, and Lauren J O’Donnell. 2023. A deep network for explainable prediction of non-imaging phenotypes using anatomical multi-view data. In International Workshop on Computation...
2023
-
[66]
Haibin Wu, Xuanjun Chen, Yi-Cheng Lin, Kai-wei Chang, Ho-Lam Chung, Alexander H Liu, and Hung-yi Lee. 2024. Towards audio language modeling--an overview. arXiv preprint arXiv:2402.13236
2024 arXiv
-
[67]
Yongchao Wu, Aron Henriksson, Martin Duneld, and Jalal Nouri. 2023. Towards improving the reliability and transparency of chatgpt for educational question answering. In European Conference on Technology Enhanced Learning
2023
-
[68]
Binzhu Xie, Sicheng Zhang, Zitang Zhou, Bo Li, Yuanhan Zhang, Jack Hessel, Jingkang Yang, and Ziwei Liu. 2024. Funqa: Towards surprising video comprehension. In European Conference on Computer Vision
2024
-
[69]
Pinci Yang, Xin Wang, Xuguang Duan, Hong Chen, Runze Hou, Cong Jin, and Wenwu Zhu. 2022. Avqa: A dataset for audio-visual question answering on videos. In International Conference on Multimedia
2022
-
[70]
Qian Yang, Jin Xu, Wenrui Liu, Yunfei Chu, Ziyue Jiang, Xiaohuan Zhou, Yichong Leng, Yuanjun Lv, Zhou Zhao, Chang Zhou, et al. 2024. Air-bench: Benchmarking large audio-language models via generative comprehension. arXiv preprint arXiv:2402.07729
2024 arXiv
-
[71]
Jiawei Yao, Qi Qian, and Juhua Hu. 2024 a . Customized multiple clustering via multi-modal subspace proxy learning. arXiv preprint arXiv:2411.03978
2024 arXiv
-
[72]
Jiawei Yao, Qi Qian, and Juhua Hu. 2024 b . Multi-modal proxy learning towards personalized visual multiple clustering. In Conference on Computer Vision and Pattern Recognition
2024
-
[73]
Qilang Ye, Zitong Yu, and Xin Liu. 2024. Answering diverse questions via text attached with key audio-visual clues. arXiv preprint arXiv:2403.06679
2024 arXiv
-
[74]
Wenhao You, Xingjian Diao, Chunhui Zhang, Keyi Kong, Weiyi Wu, Zhongyu Ouyang, Chiyu Ma, Tingxuan Wu, Noah Wei, Zong Ke, Ming Cheng, Soroush Vosoughi, and Jiang Gui. 2025. Music's multimodal complexity in avqa: Why we need more than general multimodal llms. arXiv preprint arXi...
2025 arXiv
-
[75]
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. 2024. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In Conference on Computer Vision and Pat...
2024
-
[76]
Heeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee, and Gunhee Kim. 2021. Pano-avqa: Grounded audio-visual question answering on 360deg videos. In International Conference on Computer Vision
2021
-
[77]
Chunhui Zhang, Yiren Jian, Zhongyu Ouyang, and Soroush Vosoughi. 2025. Pretrained image-text models are secretly video captioners. arXiv preprint arXiv:2502.13363
2025 arXiv
-
[78]
Jingyi Zhang, Jiaxing Huang, Sheng Jin, and Shijian Lu. 2024. Vision-language models for vision tasks: A survey. Transactions on Pattern Analysis and Machine Intelligence
2024
-
[79]
Hang Zhao, Chuang Gan, Wei-Chiu Ma, and Antonio Torralba. 2019. The sound of motions. In International Conference on Computer Vision
2019
-
[80]
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba. 2018. The sound of pixels. In European Conference on Computer Vision
2018
-
[81]
Ziyi Zhou, Ming Cheng, Xingjian Diao, Yanjun Cui, and Xiangling Li. 2024. Glumarker: A novel predictive modeling of glycemic control through digital biomarkers. In Annual International Conference of the IEEE Engineering in Medicine and Biology Society
2024
-
[82]
Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. 2021. Tat-qa: A question answering benchmark on a hybrid of tabular and textual content in finance. arXiv preprint arXiv:2105.07624
2021 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.