REVIEW 2 cited by
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Grounded Video Question Answering (Grounded VideoQA) requires aligning textual answers with explicit visual evidence. However, modern multimodal models often rely on linguistic priors and spurious correlations, resulting in poorly grounded predictions. In this work, we propose MUPA, a cooperative MUlti-Path Agentic approach that unifies video grounding, question answering, answer reflection and aggregation to tackle Grounded VideoQA. MUPA features three distinct reasoning paths on the interplay of grounding and QA agents in different chronological orders, along with a dedicated reflection agent to judge and aggregate the multi-path results to accomplish consistent QA and grounding. This design markedly improves grounding fidelity without sacrificing answer accuracy. Despite using only 2B parameters, our method outperforms all 7B-scale competitors. When scaled to 7B parameters, MUPA establishes new state-of-the-art results, with Acc@GQA of 30.3% and 47.4% on NExT-GQA and DeVE-QA respectively, demonstrating MUPA' effectiveness towards trustworthy video-language understanding. Our code is available in https://github.com/longmalongma/MUPA.
Forward citations
Cited by 2 Pith papers
-
AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding
A training-free multi-agent explore-verify system with four specialized agents and a shared evidence registry outperforms zero-shot and RL-finetuned baselines on video anomaly understanding benchmarks.
-
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
A survey paper maps how external tools are used to augment multimodal large language models across data, tasks, evaluation, and future directions.
Reference graph
Works this paper leans on
-
[1]
Revealing single frame bias for video-and-language learning,
J. Lei, T. L. Berg, and M. Bansal, “Revealing single frame bias for video-and-language learning,” arXiv preprint arXiv:2206.03428, 2022
arXiv 2022
-
[2]
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution,
P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, et al., “Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution,” arXiv preprint arXiv:2409.12191, 2024
arXiv 2024
-
[3]
Rethinking Multi-Modal Alignment in Video Question Answering from Feature and Sample Perspectives
S. Xiao, L. Chen, K. Gao, Z. Wang, Y . Yang, Z. Zhang, and J. Xiao, “Rethinking multi-modal alignment in video question answering from feature and sample perspectives,”arXiv preprint arXiv:2204.11544, 2022
work page Pith review arXiv 2022
-
[4]
Can I trust your answer? Visually grounded video question answering,
J. Xiao, A. Yao, Y . Li, and T.-S. Chua, “Can I trust your answer? Visually grounded video question answering,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024, pp. 13204–13214
2024
-
[5]
Tvqa: Localized, compositional video question answering,
Lei, Jie, Yu, Licheng, Bansal, Mohit, et al., “Tvqa: Localized, compositional video question answering,” arXiv preprint arXiv:1809.01696, 2018
arXiv 2018
-
[6]
Zero-shot video question answering via frozen bidirectional language models,
Yang, Antoine, Miech, Antoine, Sivic, Josef,et al., “Zero-shot video question answering via frozen bidirectional language models,” Advances in Neural Information Processing Systems, 2022, pp. 124–141
2022
-
[7]
Demonstrating and reducing shortcuts in vision-language representation learning,
Bleeker, Maurits, Hendriksen, Mariya, Yates, Andrew, et al., “Demonstrating and reducing shortcuts in vision-language representation learning,” arXiv preprint arXiv:2402.17510, 2024
arXiv 2024
-
[8]
Question-Answering Dense Video Events,
Qin, Hangyu, Xiao, Junbin, Yao, Angela, “Question-Answering Dense Video Events,”arXiv preprint arXiv:2409.04388, 2024
arXiv 2024
Show all 66 references
-
[9]
Grounding action descrip- tions in videos,
Regneri, Michaela, Rohrbach, Marcus, Wetzel, Dominikus, et al., “Grounding action descrip- tions in videos,” Transactions of the Association for Computational Linguistics , vol. 1, pp. 25–36, 2013
2013
-
[10]
Qvhighlights: Detecting moments and highlights in videos via natural language queries.(2021),
Lei, J, Berg, TL, Bansal, M, “Qvhighlights: Detecting moments and highlights in videos via natural language queries.(2021),” URL https://arxiv.org/abs/2107.09609, 2021
2021 arXiv
-
[11]
Videoagent: Long-form video understanding with large language model as agent,
Wang, Xiaohan, Zhang, Yuhui, Zohar, Orr,et al., “Videoagent: Long-form video understanding with large language model as agent,” in European Conference on Computer Vision, 2024, pp. 58–76
2024
-
[12]
Morevqa: Exploring modular reason- ing models for video question answering,
Min, Juhong, Buch, Shyamal, Nagrani, Arsha, et al., “Morevqa: Exploring modular reason- ing models for video question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13235–13245
2024
-
[13]
Videoagent: A memory-augmented multimodal agent for video understanding,
Fan, Yue, Ma, Xiaojian, Wu, Rujie, et al., “Videoagent: A memory-augmented multimodal agent for video understanding,” in European Conference on Computer Vision, 2024, pp. 75–92. 10
2024
-
[14]
Retrieval-based video language model for efficient long video question answering,
Xu, Jiaqi, Lan, Cuiling, Xie, Wenxuan, et al., “Retrieval-based video language model for efficient long video question answering,” arXiv preprint arXiv:2312.04931, 2023
2023
-
[15]
Hierarchical video-moment retrieval and step-captioning,
Zala, Abhay, Cho, Jaemin, Kottur, Satwik, et al., “Hierarchical video-moment retrieval and step-captioning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 23056–23065
2023
-
[16]
Tgif-qa: Toward spatio-temporal reasoning in visual question answering,
Jang, Yunseok, Song, Yale, Yu, Youngjae, Kim, Youngjin, and Kim, Gunhee, “Tgif-qa: Toward spatio-temporal reasoning in visual question answering,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 2758–2766
2017
-
[17]
Video Question Answering via Gradually Refined Attention over Appearance and Motion,
Xu, Dejing, Zhao, Zhou, Xiao, Jun, et al., “Video Question Answering via Gradually Refined Attention over Appearance and Motion,” in ACM Multimedia, 2017
2017
-
[18]
Activitynet-qa: A dataset for understanding complex web videos via question answering,
Yu, Zhou, Xu, Dejing, Yu, Jun,et al., “Activitynet-qa: A dataset for understanding complex web videos via question answering,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 33, no. 01, pp. 9127–9134, 2019
2019
-
[19]
Next-qa: Next phase of question- answering to explaining temporal actions,
Xiao, Junbin, Shang, Xindi, Yao, Angela, and Chua, Tat-Seng, “Next-qa: Next phase of question- answering to explaining temporal actions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 9777–9786
2021
-
[20]
VideoMind: A Chain-of-LoRA Agent for Long Video Reasoning,
Liu, Ye, Lin, Kevin Qinghong, Chen, Chang Wen, and Shou, Mike Zheng, “VideoMind: A Chain-of-LoRA Agent for Long Video Reasoning,” in arXiv preprint arXiv:2503.13444, 2025
2025
-
[21]
Videoqa in the era of llms: An empirical study,
Xiao, Junbin, Huang, Nanxin, Qin, Hangyu, et al., “Videoqa in the era of llms: An empirical study,”International Journal of Computer Vision, pp. 1–24, 2025
2025
-
[22]
Grounded-videollm: Sharpening fine-grained temporal grounding in video large language models,
Wang, Haibo, Xu, Zhiyang, Cheng, Yu,et al., “Grounded-videollm: Sharpening fine-grained temporal grounding in video large language models,” arXiv preprint arXiv:2410.03290, 2024
2024 arXiv
-
[23]
Cinepile: A long video question answering dataset and benchmark,
Rawal, Ruchit, Saifullah, Khalid, Farré, Miquel, et al., “Cinepile: A long video question answering dataset and benchmark,” arXiv preprint arXiv:2405.08813, 2024
2024 arXiv
-
[24]
The surprising effectiveness of multimodal large language models for video moment retrieval,
Meinardus, Boris, Batra, Anil, Rohrbach, Anna, and Rohrbach, Marcus, “The surprising effectiveness of multimodal large language models for video moment retrieval,”arXiv preprint arXiv:2406.18113, 2024
2024
-
[25]
Neptune: The Long Orbit to Bench- marking Long Video Understanding,
Nagrani, Arsha, Zhang, Mingda, Mehran, Ramin, et al., “Neptune: The Long Orbit to Bench- marking Long Video Understanding,”arXiv preprint arXiv:2412.09582, 2024
2024 arXiv
-
[26]
Self-consistency improves chain of thought reasoning in language models,
Wang, Xuezhi, Wei, Jason, Schuurmans, Dale, Le, Quoc, Chi, Ed, Narang, Sharan, Chowdh- ery, Aakanksha, and Zhou, Denny, “Self-consistency improves chain of thought reasoning in language models,” arXiv preprint arXiv:2203.11171, 2022
2022 arXiv
-
[27]
Chain-of-thought prompting elicits reasoning in large language models,
Wei, Jason, Wang, Xuezhi, Schuurmans, Dale, et al., “Chain-of-thought prompting elicits reasoning in large language models,” Advances in Neural Information Processing Systems, vol. 35, pp. 24824–24837, 2022
2022
-
[28]
Tree of thoughts: Deliberate problem solving with large language models,
Yao, Shunyu, Yu, Dian, Zhao, Jeffrey,et al., “Tree of thoughts: Deliberate problem solving with large language models,” Advances in Neural Information Processing Systems, vol. 36, pp. 11809–11822, 2023
2023
-
[29]
Enhancing LLM Reasoning with Multi-Path Collaborative Reactive and Reflection agents,
He, Chengbo, Zou, Bochao, Li, Xin, et al., “Enhancing LLM Reasoning with Multi-Path Collaborative Reactive and Reflection agents,”arXiv preprint arXiv:2501.00430, 2024
2024 arXiv
-
[30]
React: Synergizing reasoning and acting in language models,
Yao, Shunyu, Zhao, Jeffrey, Yu, Dian, et al., “React: Synergizing reasoning and acting in language models,” in International Conference on Learning Representations (ICLR), 2023
2023
-
[31]
Reflexion: Language agents with verbal reinforcement learning,
Shinn, Noah, Cassano, Federico, Gopinath, Ashwin, et al., “Reflexion: Language agents with verbal reinforcement learning,” Advances in Neural Information Processing Systems, vol. 36, pp. 8634–8652, 2023
2023
-
[32]
Improving multi-agent debate with sparse communication topology,
Li, Yunxuan, Du, Yibing, Zhang, Jiageng, et al., “Improving multi-agent debate with sparse communication topology,”arXiv preprint arXiv:2406.11776, 2024
2024 arXiv
-
[33]
An empirical study of end-to-end video-language transformers with masked visual modeling,
Fu, Tsu-Jui, Li, Linjie, Gan, Zhe, et al., “An empirical study of end-to-end video-language transformers with masked visual modeling,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 22898–22909
2023
-
[34]
Lan- guage repository for long video understanding,
Kahatapitiya, Kumara, Ranasinghe, Kanchana, Park, Jongwoo, and Ryoo, Michael S., “Lan- guage repository for long video understanding,” arXiv preprint arXiv:2403.14622, 2024. 11
2024 arXiv
-
[35]
Streaming long video understanding with large language models,
Qian, Rui, Dong, Xiaoyi, Zhang, Pan, et al., “Streaming long video understanding with large language models,” Advances in Neural Information Processing Systems, vol. 37, pp. 119336– 119360, 2024
2024
-
[36]
A simple LLM framework for long-range video question-answering,
Zhang, Ce, Lu, Taixi, Islam, Md Mohaiminul, et al., “A simple LLM framework for long-range video question-answering,” arXiv preprint arXiv:2312.17235, 2023
2023 arXiv
-
[37]
Hawkeye: Training video-text LLMs for grounding text in videos,
Wang, Yueqian, Meng, Xiaojun, Liang, Jianxin,et al., “Hawkeye: Training video-text LLMs for grounding text in videos,” arXiv preprint arXiv:2403.10228, 2024
2024 arXiv
-
[38]
Task preference optimization: Improving multimodal large language models with vision task alignment,
Yan, Ziang, Li, Zhilin, He, Yinan,et al., “Task preference optimization: Improving multimodal large language models with vision task alignment,” arXiv preprint arXiv:2412.19326, 2024
2024 arXiv
-
[39]
TVR: A large-scale dataset for video-subtitle moment retrieval,
Lei, Jie, Yu, Licheng, Berg, Tamara L., and Bansal, Mohit, “TVR: A large-scale dataset for video-subtitle moment retrieval,” in ECCV, 2020, pp. 447–463
2020
-
[40]
UMT: Unified multi-modal transformers for joint video moment retrieval and highlight detection,
Liu, Ye, Li, Siyuan, Wu, Yang,et al., “UMT: Unified multi-modal transformers for joint video moment retrieval and highlight detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 3042–3051
2022
-
[41]
MomentDiff: Generative video moment retrieval from random to real,
Li, Pandeng, Xie, Chen-Wei, Xie, Hongtao, et al., “MomentDiff: Generative video moment retrieval from random to real,” Advances in Neural Information Processing Systems, vol. 36, pp. 65948–65966, 2023
2023
-
[42]
Query-dependent video representation for moment retrieval and highlight detection,
Moon, WonJun, Hyun, Sangeek, Park, SangUk, et al., “Query-dependent video representation for moment retrieval and highlight detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 23023–23033
2023
-
[43]
UnivTG: Towards unified video- language temporal grounding,
Lin, Kevin Qinghong, Zhang, Pengchuan, Chen, Joya, et al., “UnivTG: Towards unified video- language temporal grounding,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 2794–2804
2023
-
[44]
R2-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding,
Liu, Ye, He, Jixuan, Li, Wanhua,et al., “R2-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding,” inECCV, 2024, pp. 421–438
2024
-
[45]
Learning 2D temporal adjacent networks for moment localization with natural language,
Zhang, Songyang, Peng, Houwen, Fu, Jianlong, and Luo, Jiebo, “Learning 2D temporal adjacent networks for moment localization with natural language,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 07, pp. 12870–12877, 2020
2020
-
[46]
Span-based localizing network for natural language video localization,
Zhang, Hao, Sun, Aixin, Jing, Wei, and Zhou, Joey Tianyi, “Span-based localizing network for natural language video localization,” arXiv preprint arXiv:2004.13931, 2020
2004 arXiv
-
[47]
Dense- captioning events in videos,
Krishna, Ranjay, Hata, Kenji, Ren, Frederic, Fei-Fei, Li, and Niebles, Juan Carlos, “Dense- captioning events in videos,” inProceedings of the IEEE International Conference on Computer Vision, 2017, pp. 706–715
2017
-
[48]
Negative sample matters: A renaissance of metric learning for temporal grounding,
Wang, Zhenzhi, Wang, Limin, Wu, Tao, Li, Tianhao, and Wu, Gangshan, “Negative sample matters: A renaissance of metric learning for temporal grounding,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 3, pp. 2613–2623, 2022
2022
-
[49]
Towards generalisable video moment retrieval: Visual-dynamic injection to image-text pre-training,
Luo, Dezhao, Huang, Jiabo, Gong, Shaogang, Jin, Hailin, and Liu, Yang, “Towards generalisable video moment retrieval: Visual-dynamic injection to image-text pre-training,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 23045– 23055
2023
-
[50]
VideoChat: Chat-centric video understanding,
Li, KunChang, He, Yinan, Wang, Yi, et al., “VideoChat: Chat-centric video understanding,” arXiv preprint arXiv:2305.06355, 2023
2023 arXiv
-
[51]
Video-LLaMA: An instruction-tuned audio-visual language model for video understanding,
Zhang, Hang, Li, Xin, and Bing, Lidong, “Video-LLaMA: An instruction-tuned audio-visual language model for video understanding,” arXiv preprint arXiv:2306.02858, 2023
2023 arXiv
-
[52]
Video- ChatGPT: Towards detailed video understanding via large vision and language models,
Maaz, Muhammad, Rasheed, Hanoona, Khan, Salman, and Khan, Fahad Shahbaz, “Video- ChatGPT: Towards detailed video understanding via large vision and language models,”arXiv preprint arXiv:2306.05424, 2023
2023 arXiv
-
[53]
V ALLEY: Video assistant with large language model enhanced ability,
Luo, Ruipu, Zhao, Ziwang, Yang, Min, et al., “V ALLEY: Video assistant with large language model enhanced ability,”arXiv preprint arXiv:2306.07207, 2023
2023 arXiv
-
[54]
ChatVTG: Video temporal grounding via chat with video dialogue large language models,
Qu, Mengxue, Chen, Xiaodong, Liu, Wu, Li, Alicia, and Zhao, Yao, “ChatVTG: Video temporal grounding via chat with video dialogue large language models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 1847–1856. 12
2024
-
[55]
Momentor: Advancing video large language model with fine-grained temporal reasoning,
Qian, Long, Li, Juncheng, Wu, Yu,et al., “Momentor: Advancing video large language model with fine-grained temporal reasoning,” arXiv preprint arXiv:2402.11435, 2024
2024 arXiv
-
[56]
ET Bench: Towards open-ended event-level video-language understanding,
Liu, Ye, Ma, Zongyang, Qi, Zhongang, Wu, Yang, Shan, Ying, and Chen, Chang W, “ET Bench: Towards open-ended event-level video-language understanding,”Advances in Neural Information Processing Systems, vol. 37, pp. 32076–32110, 2024
2024
-
[57]
LITA: Language instructed temporal-localization assistant,
Huang, De-An, Liao, Shijia, Radhakrishnan, Subhashree, et al., “LITA: Language instructed temporal-localization assistant,” in European Conference on Computer Vision, 2024, pp. 202– 218
2024
-
[58]
RexTime: A benchmark suite for reasoning- across-time in videos,
Chen, Jr-Jen, Liao, Yu-Chien, Lin, Hsi-Che,et al., “RexTime: A benchmark suite for reasoning- across-time in videos,” Advances in Neural Information Processing Systems , vol. 37, pp. 28662–28673, 2024
2024
-
[59]
VTimeLLM: Empower LLM to grasp video moments,
Huang, Bin, Wang, Xin, Chen, Hong, Song, Zihan, and Zhu, Wenwu, “VTimeLLM: Empower LLM to grasp video moments,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 14271–14280
2024
-
[60]
TimeChat: A time-sensitive mul- timodal large language model for long video understanding,
Ren, Shuhuai, Yao, Linli, Li, Shicheng, Sun, Xu, and Hou, Lu, “TimeChat: A time-sensitive mul- timodal large language model for long video understanding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 14313–14323
2024
-
[61]
LoRA: Low-rank adaptation of large language models,
Hu, Edward J, Shen, Yelong, Wallis, Phillip, et al., “LoRA: Low-rank adaptation of large language models,” International Conference on Learning Representations (ICLR), vol. 1, no. 2, pp. 3, 2022
2022
-
[62]
Self-chained image-language model for video localization and question answering,
Yu, Shoubin, Cho, Jaemin, Yadav, Prateek, and Bansal, Mohit, “Self-chained image-language model for video localization and question answering,” Advances in Neural Information Process- ing Systems, vol. 36, pp. 76749–76771, 2023
2023
-
[63]
GPT-4o System Card,
Hurst, Aaron, Lerer, Adam, Goucher, Adam P,et al., “GPT-4o System Card,”arXiv preprint arXiv:2410.21276, 2024
2024 arXiv
-
[64]
Cosmo: Contrastive streamlined multimodal model with interleaved pre-training,
Wang, Alex Jinpeng, Li, Linjie, Lin, Kevin Qinghong,et al., “Cosmo: Contrastive streamlined multimodal model with interleaved pre-training,” arXiv preprint arXiv:2401.00849, 2024
2024 arXiv
-
[65]
Reinforcing video reasoning with focused thinking,
Dang, Jisheng, Wu, Jingze, Wang, Teng, et al., “Reinforcing video reasoning with focused thinking,” arXiv preprint arXiv:2505.24718, 2025
2025 arXiv
-
[66]
SynPO: Synergizing descriptiveness and preference optimization for video detailed captioning,
Dang, Jisheng, Zhang, Yizhou, Ye, Hao et al. , “SynPO: Synergizing descriptiveness and preference optimization for video detailed captioning,” arXiv preprint arXiv:2506.00835, 2025. 13 A Discussion on Entropy-based Measurement In this document, we provide more descriptions of ...
2025
Discussion (0). Sign in to comment.