REVIEW 4 major objections 3 minor 3 cited by
A Survey on Video Temporal Grounding with Multimodal Large Language Model
T0 review · 4 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This survey argues that video temporal grounding methods built on multimodal large language models are surpassing traditional fine-tuned approaches, and organizes the field with a three-dimensional taxonomy.
desk verdict Useful-looking survey of VTG-MLLMs; only the abstract is legible, so the core empirical claim is unsupported in the artifact we have. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key machinery is the three-dimensional taxonomy itself: (1) functional roles of MLLMs, highlighting architectural significance; (2) training paradigms for temporal reasoning and task adaptation; and (3) video feature processing techniques that determine spatiotemporal representation. This taxonomy is what allows the survey to compare methods systematically, summarize empirical findings, and point to limitations and research directions.
What would settle it
A concrete finding that would undermine the central claim: a published VTG-MLLM method that cannot be placed in any of the three taxonomy axes, or a rigorous multi-domain benchmark where traditional fine-tuned methods consistently beat MLLM-based methods.
Extended reading notes
Core claim
The central claim of the survey is that MLLM-based methods have become the leading paradigm for video temporal grounding. In the paper's own terms, VTG-MLLMs 'are gradually surpassing traditional fine-tuned methods': they stay competitive on standard benchmarks and excel in zero-shot, multi-task, and multi-domain generalization. The contribution is a three-dimensional taxonomy that classifies each approach by the role the MLLM plays, the training strategy used, and the video feature processing pipeline, together with a review of benchmarks, evaluation protocols, and empirical findings.
Load-bearing premise
The load-bearing premise is that the three-axis taxonomy cleanly classifies every VTG-MLLM approach and that the reported benchmark comparisons are representative of the field.
Editorial extensions
If this is right
- Future VTG-MLLM papers can be positioned within a shared three-dimensional space, making method comparisons and reproducibility clearer.
- The empirical trend implies that new VTG research should build on MLLMs rather than fine-tuned baselines, especially when targeting zero-shot or multi-domain deployment.
- Benchmark design should include zero-shot, multi-task, and multi-domain settings to capture the claimed generalization advantage.
- The survey's identified limitations and proposed directions provide a concrete agenda for improving temporal reasoning and task adaptation.
Reading between the lines
- The three-axis taxonomy may extend to neighboring video-language tasks such as video captioning or spatio-temporal grounding, since the axes are not specific to temporal localization.
- The claim that MLLM methods surpass fine-tuned methods may depend on benchmark and domain choice; a standardized, broader evaluation would test how universal the advantage is.
- In practice the three axes may overlap (training paradigms can influence feature processing), so the taxonomy is best read as complementary lenses rather than a strict partition.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of video temporal grounding with multimodal large language models (VTG-MLLMs). The abstract claims that VTG-MLLMs are gradually surpassing traditional fine-tuned methods, both in competitive performance and in generalization across zero-shot, multi-task, and multi-domain settings. The survey is organized around a three-dimensional taxonomy: functional roles of MLLMs, training paradigms, and video feature processing techniques. It also promises to cover benchmark datasets, evaluation protocols, empirical findings, limitations, and future research directions, with an accompanying GitHub repository. The supplied full text, however, is almost entirely composed of Unicode replacement characters, making the body and tables unreadable. Consequently, the survey's substantive content—its taxonomies, classifications, benchmark comparisons, and critical discussion—cannot be assessed from the provided artifact; only the abstract and the presence of an unreadable appendix/table section are available.
Significance. If the claims in the abstract were fully supported by the body, the survey would make a useful contribution by providing an organizing three-axis taxonomy for a rapidly growing but scattered literature, and by synthesizing empirical evidence on whether MLLM-based approaches truly outperform traditional fine-tuned methods. However, the significance of the empirical claim depends entirely on the completeness, accuracy, and protocol-matching of benchmark comparisons, which are not visible in the supplied full text. The taxonomy's value similarly depends on whether the three axes are disjoint and jointly exhaustive, a property that is not established in the abstract and cannot be checked. No machine-checked proofs, reproducible code, or parameter-free derivations are present. The manuscript's potential significance is therefore real but currently unverifiable from the artifact provided.
major comments (4)
- [Abstract vs. Full Text] The central empirical claim, that 'VTG approaches based on MLLMs are gradually surpassing traditional fine-tuned methods' and generalize better across zero-shot, multi-task, and multi-domain settings, is load-bearing. In the supplied full text, the body after the abstract is overwhelmingly composed of U+FFFD replacement characters, and the tables on the later pages are unreadable. No benchmark numbers, evaluation protocols, error bars, or controlled comparisons are visible. This is a verification gap rather than an internal inconsistency, but it means the survey's main empirical conclusion is unsupported in the provided artifact. Please provide a readable manuscript with the actual text, tables, and references so the evidence can be assessed.
- [Three-dimensional taxonomy] The abstract introduces a three-axis taxonomy—functional roles of MLLMs, training paradigms, and video feature processing—as the organizing structure. The taxonomy is only stated, not demonstrated to be exhaustive or disjoint. In particular, 'functional role' and 'training paradigm' may overlap: a method's functional role is often expressed through its training objective or adapter design. Without explicit definitions of category boundaries and a classification of representative methods, a reader cannot judge whether the taxonomy systematically covers all current VTG-MLLM work or whether some methods fall between categories. This is a load-bearing organizational claim that requires concrete support in the survey body.
- [Benchmarks section] The abstract promises discussion of benchmark datasets, evaluation protocols, and empirical findings. In the supplied full text, the only readable section appears to be a table-like block of replacement characters, which cannot be interpreted. There is no verifiable presentation of datasets, metrics, or results. Since the survey's contribution includes 'summarize empirical findings,' the absence of legible benchmark analysis prevents any validation of the claimed performance trends. Please provide the benchmark tables in machine-readable or otherwise legible form, with protocol details and citations.
- [Limitations and future directions] The abstract states that the survey 'identifies existing limitations and proposes promising research directions.' These sections are also unreadable in the supplied artifact. A survey's critical value depends on its assessment of limitations (e.g., evaluation biases, benchmark saturation, generalization gaps). Without access to these sections, the completeness and even-handedness of the survey cannot be judged.
minor comments (3)
- [Full text encoding] The supplied PDF/HTML appears to have a text-encoding failure: almost all body text is replaced by U+FFFD replacement characters. This must be fixed in the submission; a clean, readable version is essential for review and for reader use.
- [Repository link] The abstract refers to https://github.com/ki-lw/Awesome-MLLMs-for-Video-Temporal-Grounding for additional resources. The repository may be useful, but it is not a substitute for the survey's own text and tables. Please ensure the survey body stands alone independently of the repository.
- [References] The provided text does not show a reference list. A survey must include complete citations for all discussed methods and benchmarks; these need to be legible and correct.
Circularity Check
No circularity identified: this is a survey with no derivation chain; the main empirical claim is unverifiable from the garbled body but is not shown to reduce to its own inputs.
full rationale
The paper is a survey of video temporal grounding with multimodal large language models. Its contribution is a three-dimensional taxonomy (functional roles of MLLMs, training paradigms, video feature processing) and an empirical synthesis claiming that VTG-MLLM methods are 'gradually surpassing traditional fine-tuned methods' and generalize better in zero-shot, multi-task, and multi-domain settings. A survey does not ordinarily present a derivation chain of the kind that could be circular: there are no equations, no fitted parameters renamed as predictions, and no formal claims that reduce to their definitions. The abstract's empirical claim depends on benchmark evidence, and the supplied full text is largely unreadable U+FFFD replacement characters, so those benchmark tables cannot be audited. However, unverifiability is not circularity. The taxonomy's axes may be overlapping (a method's functional role is often expressed through its training paradigm), which is a classification-quality concern rather than a circular-reasoning concern. No load-bearing self-citation can be identified from the provided artifact. Under the hard rules requiring a quotable specific reduction, no circular step can be exhibited. The honest finding is therefore no significant circularity, score 0.
Assumptions & free parameters
assumptions (2)
- domain assumption The three-axis taxonomy (functional roles, training paradigms, video feature processing) is exhaustive and categories are disjoint.
- domain assumption The empirical claim that VTG-MLLMs surpass traditional fine-tuned methods is faithfully reported from the cited benchmarks.
Cite this review
Pith. "Pith review of A Survey on Video Temporal Grounding with Multimodal Large Language Model." pith.science (2026). https://pith.science/paper/DAUTZFKX
@misc{pith2026250810922,
author = {Pith},
title = {Pith review of: A Survey on Video Temporal Grounding with Multimodal Large Language Model},
year = {2026},
howpublished = {\url{https://pith.science/paper/DAUTZFKX}},
note = {Machine review of arXiv:2508.10922}
}
read the original abstract
The recent advancement in video temporal grounding (VTG) has significantly enhanced fine-grained video understanding, primarily driven by multimodal large language models (MLLMs). With superior multimodal comprehension and reasoning abilities, VTG approaches based on MLLMs (VTG-MLLMs) are gradually surpassing traditional fine-tuned methods. They not only achieve competitive performance but also excel in generalization across zero-shot, multi-task, and multi-domain settings. Despite extensive surveys on general video-language understanding, comprehensive reviews specifically addressing VTG-MLLMs remain scarce. To fill this gap, this survey systematically examines current research on VTG-MLLMs through a three-dimensional taxonomy: 1) the functional roles of MLLMs, highlighting their architectural significance; 2) training paradigms, analyzing strategies for temporal reasoning and task adaptation; and 3) video feature processing techniques, which determine spatiotemporal representation effectiveness. We further discuss benchmark datasets, evaluation protocols, and summarize empirical findings. Finally, we identify existing limitations and propose promising research directions. For additional resources and details, readers are encouraged to visit our repository at https://github.com/ki-lw/Awesome-MLLMs-for-Video-Temporal-Grounding.
Forward citations
Cited by 3 Pith papers
-
NEST: Narrative Event Structures in Time for Long Video Understanding
NEST is a new benchmark dataset for narrative event structures in long videos, with baselines reporting ETD below 8%, EL under 6%, EAE below 11%, and ERE at 35-44% F1.
-
Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition
Hour-long video temporal grounding is a search problem, shown by a new benchmark where all Video-LLMs collapse, frame retrieval outperforms them, 85% of failures are search-related, and a retrieve-then-ground hybrid i...
-
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
Selecting the visual input that minimizes an MLLM's output entropy (or maximizes its yes/no confidence) improves fine-grained visual search, long-video QA, and temporal grounding without any training.
Reference graph
Works this paper leans on
-
[1]
11em plus .33em minus .07em 4000 4000 100 4000 4000 500 `\.=1000 = #1 \@IEEEnotcompsoconly \@IEEEcompsoconly #1 * [1] 0pt [0pt][0pt] #1 * [1] 0pt [0pt][0pt] #1 * \| ** #1 \@IEEEauthorblockNstyle \@IEEEcompsocnotconfonly \@IEEEauthorblockAstyle \@IEEEcompsocnotconfonly \@IEEEcompsocconfonly \@IEEEauthordefaulttextstyle \@IEEEcompsocnotconfonly \@IEEEauthor...
-
[2]
Y. Wang, K. Li, Y. Li, Y. He, B. Huang, Z. Zhao, H. Zhang, J. Xu, Y. Liu, Z. Wang et al., ``Internvideo: General video foundation models via generative and discriminative learning,'' arXiv:2212.03191, 2022
arXiv 2022
-
[3]
Y. Wang, K. Li, X. Li, J. Yu, Y. He, G. Chen, B. Pei, R. Zheng, Z. Wang, Y. Shi et al., ``Internvideo2: Scaling foundation models for multimodal video understanding,'' in ECCV, 2024, pp. 396--416
2024
-
[4]
M. Wang, J. Xing, J. Mei, Y. Liu, and Y. Jiang, ``Actionclip: Adapting language-image pretrained models for video action recognition,'' TNNLS, vol. 36, no. 1, pp. 625--637, 2023
2023
-
[5]
A. Yang, A. Nagrani, P. H. Seo, A. Miech, J. Pont-Tuset, I. Laptev, J. Sivic, and C. Schmid, ``Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,'' in CVPR, 2023, pp. 10\,714--10\,726
2023
-
[6]
Regneri, M
M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal, ``Grounding action descriptions in videos,'' TACL, vol. 1, pp. 25--36, 2013
2013
-
[7]
Anne Hendricks, O
L. Anne Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell, ``Localizing moments in video with natural language,'' in ICCV, 2017, pp. 5803--5812
2017
-
[8]
Krishna, K
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. Carlos Niebles, ``Dense-captioning events in videos,'' in ICCV, 2017, pp. 706--715
2017
Show all 164 references
-
[9]
T. Wang, R. Zhang, Z. Lu, F. Zheng, R. Cheng, and P. Luo, ``End-to-end dense video captioning with parallel decoding,'' in ICCV, 2021, pp. 6847--6857
2021
-
[10]
Y. Song, J. Vallmitjana, A. Stent, and A. Jaimes, ``Tvsum: Summarizing web videos using titles,'' in CVPR, 2015, pp. 5179--5187
2015
-
[11]
Xiong, Y
B. Xiong, Y. Kalantidis, D. Ghadiyaram, and K. Grauman, ``Less is more: Learning highlight detection from video duration,'' in CVPR, 2019, pp. 1258--1267
2019
-
[12]
J. Xiao, A. Yao, Y. Li, and T.-S. Chua, ``Can i trust your answer? visually grounded video question answering,'' in CVPR, 2024, pp. 13\,204--13\,214
2024
-
[13]
Chen, Y.-C
J.-J. Chen, Y.-C. Liao, H.-C. Lin, Y.-C. Yu, Y.-C. Chen, and F. Wang, ``Rextime: A benchmark suite for reasoning-across-time in videos,'' in NeurIPS, 2024, pp. 28\,662--28\,673
2024
-
[14]
Zhang, X
D. Zhang, X. Dai, X. Wang, Y.-F. Wang, and L. S. Davis, ``Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment,'' in CVPR, 2019, pp. 1247--1257
2019
-
[15]
W. Moon, S. Hyun, S. Park, D. Park, and J.-P. Heo, ``Query-dependent video representation for moment retrieval and highlight detection,'' in CVPR, 2023, pp. 23\,023--23\,033
2023
-
[16]
M. Ma, S. Yoon, J. Kim, Y. Lee, S. Kang, and C. D. Yoo, ``Vlanet: Video-language alignment network for weakly-supervised video moment retrieval,'' in ECCV, 2020, pp. 156--171
2020
-
[17]
Zhang, Y
M. Zhang, Y. Yang, X. Chen, Y. Ji, X. Xu, J. Li, and H. T. Shen, ``Multi-stage aggregated transformer network for temporal language localization in videos,'' in CVPR, 2021, pp. 12\,669--12\,678
2021
-
[18]
Y. Liu, S. Li, Y. Wu, C.-W. Chen, Y. Shan, and X. Qie, ``Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection,'' in CVPR, 2022, pp. 3042--3051
2022
-
[19]
Y. Yuan, T. Mei, and W. Zhu, ``To find where you talk: Temporal sentence localization in video with attention based location regression,'' in AAAI, 2019, pp. 9159--9166
2019
-
[20]
Zhang, S
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin et al., ``Opt: Open pre-trained transformer language models,'' arXiv:2205.01068, 2022
2022 arXiv
-
[21]
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma et al., ``Scaling instruction-finetuned language models,'' JMLR, vol. 25, no. 70, pp. 1--53, 2024
2024
-
[22]
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi et al., ``Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,'' arXiv:2501.12948, 2025
2025 arXiv
-
[23]
H. Liu, C. Li, Q. Wu, and Y. J. Lee, ``Visual instruction tuning,'' in NeurIPS, 2023, pp. 34\,892--34\,916
2023
-
[24]
W. Dai, J. Li, D. LI, A. Tiong, J. Zhao, W. Wang, B. Li, P. N. Fung, and S. Hoi, ``Instructblip: Towards general-purpose vision-language models with instruction tuning,'' in NeurIPS, 2023, pp. 49\,250--49\,267
2023
-
[25]
Cheng, S
Z. Cheng, S. Leng, H. Zhang, Y. Xin, X. Li, G. Chen, Y. Zhu, W. Zhang, Z. Luo, D. Zhao, and L. Bing, ``Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms,'' arXiv:2406.07476, 2024
2024 arXiv
-
[26]
L. Xu, Y. Zhao, D. Zhou, Z. Lin, S. K. Ng, and J. Feng, ``Pllava: Parameter-free llava extension from images to videos for video dense captioning,'' arXiv:2404.16994, 2024
2024 arXiv
-
[27]
Huang, X
B. Huang, X. Wang, H. Chen, Z. Song, and W. Zhu, ``Vtimellm: Empower llm to grasp video moments,'' in CVPR, 2024, pp. 14\,271--14\,280
2024
-
[28]
S. Ren, L. Yao, S. Li, X. Sun, and L. Hou, ``Timechat: A time-sensitive multimodal large language model for long video understanding,'' in CVPR, 2024, pp. 14\,313--14\,323
2024
-
[29]
L. Qian, J. Li, Y. Wu, Y. Ye, H. Fei, T.-S. Chua, Y. Zhuang, and S. Tang, ``Momentor: Advancing video large language model with fine-grained temporal reasoning,'' in ICML, 2024, pp. 41\,340--41\,356
2024
-
[30]
M. Qu, X. Chen, W. Liu, A. Li, and Y. Zhao, ``Chatvtg: Video temporal grounding via chat with video dialogue large language models,'' in CVPR, 2024, pp. 1847--1856
2024
-
[31]
H. Qin, J. Xiao, and A. Yao, ``Question-answering dense video events,'' arXiv:2409.04388, 2024
2024 arXiv
-
[32]
Y. Guo, J. Liu, M. Li, X. Tang, Q. Liu, and X. Chen, ``Trace: Temporal grounding video llm via causal event modeling,'' in ICLR, 2025
2025
-
[33]
Y. Wang, X. Meng, J. Liang, Y. Wang, Q. Liu, and D. Zhao, ``Hawkeye: Training video-text llms for grounding text in videos,'' arXiv:2403.10228, 2024
2024 arXiv
-
[34]
X. Zeng, K. Li, C. Wang, X. Li, T. Jiang, Z. Yan, S. Li, Y. Shi, Z. Yue, Y. Wang et al., ``Timesuite: Improving mllms for long video understanding via grounded tuning,'' in ICLR, 2025
2025
-
[35]
H. Wang, Z. Xu, Y. Cheng, S. Diao, Y. Zhou, Y. Cao, Q. Wang, W. Ge, and L. Huang, ``Grounded-videollm: Sharpening fine-grained temporal grounding in video large language models,'' arXiv:2410.03290, 2024
2024 arXiv
-
[36]
Y. Liu, K. Q. Lin, C. W. Chen, and M. Z. Shou, ``Videomind: A chain-of-lora agent for long video reasoning,'' arXiv:2503.13444, 2025
2025
-
[37]
C. Li, Z. Gan, Z. Yang, J. Yang, L. Li, L. Wang, J. Gao et al., ``Multimodal foundation models: From specialists to general-purpose assistants,'' arXiv:2309.10020, 2023
2023 arXiv
-
[38]
Y. Zhu, X. Li, C. Liu, M. Zolfaghari, Y. Xiong, C. Wu, Z. Zhang, J. Tighe, R. Manmatha, and M. Li, ``A comprehensive study of deep video action recognition,'' arXiv:2012.06567, 2020
2012 arXiv
-
[39]
Abdar, M
M. Abdar, M. Kollati, S. Kuraparthi, F. Pourpanah, D. McDuff, M. Ghavamzadeh, S. Yan, A. Mohamed, A. Khosravi, E. Cambria et al., ``A review of deep learning for video captioning,'' TPAMI, pp. 1--20, 2024
2024
-
[40]
Z. Xing, Q. Feng, H. Chen, Q. Dai, H. Hu, H. Xu, Z. Wu, and Y.-G. Jiang, ``A survey on video diffusion models,'' ACM Computing Surveys, vol. 57, no. 2, pp. 1--42, 2024
2024
-
[41]
Zhang, J
J. Zhang, J. Huang, S. Jin, and S. Lu, ``Vision-language models for vision tasks: A survey,'' TPAMI, vol. 46, no. 8, pp. 5625--5644, 2024
2024
-
[42]
X. Liu, X. Nie, Z. Tan, J. Guo, and Y. Yin, ``A survey on natural language video localization,'' arXiv:2104.00234, 2021
2021 arXiv
-
[43]
Y. Yang, Z. Li, and G. Zeng, ``A survey of temporal activity localization via language in untrimmed videos,'' in ICCST, 2020, pp. 596--601
2020
-
[44]
X. Lan, Y. Yuan, X. Wang, Z. Wang, and W. Zhu, ``A survey on temporal sentence grounding in videos,'' TOMCCAP, vol. 19, no. 2, pp. 1--33, 2023
2023
-
[45]
Zhang, A
H. Zhang, A. Sun, W. Jing, and J. T. Zhou, ``Temporal sentence grounding in videos: A survey and future directions,'' TPAMI, vol. 45, no. 8, pp. 10\,443--10\,465, 2023
2023
-
[46]
Zhang, H
S. Zhang, H. Peng, J. Fu, and J. Luo, ``Learning 2d temporal adjacent networks for moment localization with natural language,'' in AAAI, 2020, pp. 12\,870--12\,877
2020
-
[47]
K. Q. Lin, P. Zhang, J. Chen, S. Pramanick, D. Gao, A. J. Wang, R. Yan, and M. Z. Shou, ``Univtg: Towards unified video-language temporal grounding,'' in ICCV, 2023, pp. 2794--2804
2023
-
[48]
W. Yang, T. Zhang, Y. Zhang, and F. Wu, ``Local correspondence network for weakly supervised temporal sentence grounding,'' TIP, vol. 30, pp. 3252--3262, 2021
2021
-
[49]
D. Liu, X. Qu, J. Dong, P. Zhou, Y. Cheng, W. Wei, Z. Xu, and Y. Xie, ``Context-aware biaffine localizing network for temporal sentence grounding,'' in CVPR, 2021, pp. 11\,235--11\,244
2021
-
[50]
S. Xiao, L. Chen, S. Zhang, W. Ji, J. Shao, L. Ye, and J. Xiao, ``Boundary proposal network for two-stage natural language video localization,'' in AAAI, 2021, pp. 2986--2994
2021
-
[51]
Zhang, A
H. Zhang, A. Sun, W. Jing, L. Zhen, J. T. Zhou, and R. S. M. Goh, ``Natural language video localization: A revisit in span-based question answering framework,'' TPAMI, vol. 44, no. 8, pp. 4252--4266, 2021
2021
-
[52]
Y. Liu, Z. Ma, Z. Qi, Y. Wu, Y. Shan, and C. W. Chen, ``Et bench: Towards open-ended event-level video-language understanding,'' in NeurIPS, 2024, pp. 32\,076--32\,110
2024
-
[53]
L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong, ``End-to-end dense video captioning with masked transformer,'' in CVPR, 2018, pp. 8739--8748
2018
-
[54]
Iashin and E
V. Iashin and E. Rahtu, ``Multi-modal dense video captioning,'' in CVPRW, 2020, pp. 958--959
2020
-
[55]
J. Lei, T. L. Berg, and M. Bansal, ``Detecting moments and highlights in videos via natural language queries,'' in NeurIPS, 2021, pp. 11\,846--11\,858
2021
-
[56]
G. Chen, Y. Liu, Y. Huang, Y. He, B. Pei, J. Xu, Y. Wang, T. Lu, and L. Wang, ``Cg-bench: Clue-grounded question answering benchmark for long video understanding,'' arXiv:2412.12075, 2024
2024 arXiv
-
[57]
H. Liu, X. Ma, C. Zhong, Y. Zhang, and W. Lin, ``Timecraft: Navigate weakly-supervised temporal grounded video question answering via bi-directional reasoning,'' in ECCV, 2025, pp. 92--107
2025
-
[58]
Zhang, X
H. Zhang, X. Li, and L. Bing, ``Video-llama: An instruction-tuned audio-visual language model for video understanding,'' in EMNLP, 2023, pp. 543--553
2023
-
[59]
B. Lin, Y. Ye, B. Zhu, J. Cui, M. Ning, P. Jin, and L. Yuan, ``Video-llava: Learning united visual representation by alignment before projection,'' arXiv:2311.10122, 2023
2023 arXiv
-
[60]
K. Li, Y. He, Y. Wang, Y. Li, W. Wang, P. Luo, Y. Wang, L. Wang, and Y. Qiao, ``Videochat: Chat-centric video understanding,'' arXiv:2305.06355, 2023
2023 arXiv
-
[61]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., ``An image is worth 16x16 words: Transformers for image recognition at scale,'' in ICLR, 2020
2020
-
[62]
Q. Sun, Y. Fang, L. Wu, X. Wang, and Y. Cao, ``Eva-clip: Improved training techniques for clip at scale,'' arXiv:2303.15389, 2023
2023 arXiv
-
[63]
Y. Fang, W. Wang, B. Xie, Q. Sun, L. Wu, X. Wang, T. Huang, X. Wang, and Y. Cao, ``Eva: Exploring the limits of masked visual representation learning at scale,'' in CVPR, 2023, pp. 19\,358--19\,369
2023
-
[64]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., ``Learning transferable visual models from natural language supervision,'' in ICML, 2021, pp. 8748--8763
2021
-
[65]
Feichtenhofer, Y
C. Feichtenhofer, Y. Li, K. He et al., ``Masked autoencoders as spatiotemporal learners,'' in NeurIPS, 2022, pp. 35\,946--35\,958
2022
-
[66]
Z. Tong, Y. Song, J. Wang, and L. Wang, ``Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training,'' in NeurIPS, 2022, pp. 10\,078--10\,093
2022
-
[67]
L. Wang, B. Huang, Z. Zhao, Z. Tong, Y. He, Y. Wang, Y. Wang, and Y. Qiao, ``Videomae v2: Scaling video masked autoencoders with dual masking,'' in CVPR, 2023, pp. 14\,549--14\,560
2023
-
[68]
J. Wang, Y. Ge, R. Yan, Y. Ge, K. Q. Lin, S. Tsutsui, X. Lin, G. Cai, J. Wu, Y. Shan et al., ``All in one: Exploring unified video-language pre-training,'' in CVPR, 2023, pp. 6598--6608
2023
-
[69]
G. Sun, W. Yu, C. Tang, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, Y. Wang, and C. Zhang, ``video-salmonn: Speech-enhanced audio-visual large language models,'' in ICML, 2024, pp. 47\,198--47\,217
2024
-
[70]
S. Azad, V. Vineet, and Y. S. Rawat, ``Hierarq: Task-aware hierarchical q-former for enhanced video understanding,'' arXiv:2503.08585, 2025
2025 arXiv
-
[71]
Alayrac, J
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al., ``Flamingo: a visual language model for few-shot learning,'' in NeurIPS, 2022, pp. 23\,716--23\,736
2022
-
[72]
J. Li, D. Li, S. Savarese, and S. Hoi, ``Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,'' in ICML, 2023, pp. 19\,730--19\,742
2023
-
[73]
Q. Ye, H. Xu, G. Xu, J. Ye, M. Yan, Y. Zhou, J. Wang, A. Hu, P. Shi, Y. Shi et al., ``mplug-owl: Modularization empowers large language models with multimodality,'' arXiv:2304.14178, 2023
2023 arXiv
-
[74]
Huang, L
S. Huang, L. Dong, W. Wang, Y. Hao, S. Singhal, S. Ma, T. Lv, L. Cui, O. K. Mohammed, B. Patra et al., ``Language is not all you need: Aligning perception with language models,'' in NeurIPS, 2023, pp. 72\,096--72\,109
2023
-
[75]
H. Liu, C. Li, Y. Li, and Y. J. Lee, ``Improved baselines with visual instruction tuning,'' in CVPR, 2024, pp. 26\,296--26\,306
2024
-
[76]
F. Li, R. Zhang, H. Zhang, Y. Zhang, B. Li, W. Li, Z. Ma, and C. Li, ``Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models,'' arXiv:2407.07895, 2024
2024 arXiv
-
[77]
J. Ye, H. Xu, H. Liu, A. Hu, M. Yan, Q. Qian, J. Zhang, F. Huang, and J. Zhou, ``mplug-owl3: Towards long image-sequence understanding in multi-modal large language models,'' arXiv:2408.04840, 2024
2024 arXiv
-
[78]
K. Chen, D. Shen, H. Zhong, H. Zhong, K. Xia, D. Xu, W. Yuan, Y. Hu, B. Wen, T. Zhang et al., ``Evlm: An efficient vision-language model for visual understanding,'' arXiv:2407.14177, 2024
2024 arXiv
-
[79]
M. Shi, S. Wang, C.-Y. Chen, J. Jain, K. Wang, J. Xiong, G. Liu, Z. Yu, and H. Shi, ``Slow-fast architecture for video multi-modal large language models,'' arXiv:2504.01328, 2025
2025 arXiv
-
[80]
Miech, D
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic, ``Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,'' in ICCV, 2019, pp. 2630--2640
2019
-
[81]
M. Bain, A. Nagrani, G. Varol, and A. Zisserman, ``Frozen in time: A joint video and image encoder for end-to-end retrieval,'' in ICCV, 2021, pp. 1728--1738
2021
-
[82]
Schuhmann, R
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman et al., ``Laion-5b: An open large-scale dataset for training next generation image-text models,'' in NeurIPS, 2022, pp. 25\,278--25\,294
2022
-
[83]
L. Li, Y. Yin, S. Li, L. Chen, P. Wang, S. Ren, M. Li, Y. Yang, J. Xu, X. Sun et al., ``M ^3 it: A large-scale dataset towards multi-modal multilingual instruction tuning,'' arXiv:2306.04387, 2023
2023 arXiv
-
[84]
Z. Yin, J. Wang, J. Cao, Z. Shi, D. Liu, M. Li, X. Huang, Z. Wang, L. Sheng, L. Bai et al., ``Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark,'' in NeurIPS, 2023, pp. 26\,650--26\,685
2023
-
[85]
H. Shao, S. Qian, H. Xiao, G. Song, Z. Zong, L. Wang, Y. Liu, and H. Li, ``Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning,'' in NeurIPS, 2024, pp. 8612--8642
2024
-
[86]
T. Gao, P. Chen, M. Zhang, C. Fu, Y. Shen, Y. Zhang, S. Zhang, X. Zheng, X. Sun, L. Cao et al., ``Cantor: Inspiring multimodal chain-of-thought of mllm,'' in ACM MM, 2024, pp. 9096--9105
2024
-
[87]
J. Wu, Z. Zhang, Y. Xia, X. Li, Z. Xia, A. Chang, T. Yu, S. Kim, R. A. Rossi, R. Zhang et al., ``Visual prompting in multimodal large language models: A survey,'' arXiv:2409.15310, 2024
2024 arXiv
-
[88]
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen et al., ``Lora: Low-rank adaptation of large language models,'' in ICLR, 2022
2022
-
[89]
R. Pan, X. Liu, S. Diao, R. Pi, J. Zhang, C. Han, and T. Zhang, ``Lisa: layerwise importance sampling for memory-efficient large language model fine-tuning,'' in NeurIPS, 2024, pp. 57\,018--57\,049
2024
-
[90]
Liu, C.-Y
S.-Y. Liu, C.-Y. Wang, H. Yin, P. Molchanov, Y.-C. F. Wang, K.-T. Cheng, and M.-H. Chen, ``Dora: Weight-decomposed low-rank adaptation,'' in ICML, 2024, pp. 32\,100--32\,121
2024
-
[91]
W. Cai, J. Huang, S. Gong, H. Jin, and Y. Liu, ``Mllm as video narrator: Mitigating modality imbalance in video moment retrieval,'' arXiv:2406.17880, 2024
2024 arXiv
-
[92]
Di and W
S. Di and W. Xie, ``Grounded question-answering in long egocentric videos,'' in CVPR, 2024, pp. 12\,934--12\,943
2024
-
[93]
Zheng, X
M. Zheng, X. Cai, Q. Chen, Y. Peng, and Y. Liu, ``Training-free video temporal grounding using large-scale pre-trained models,'' in ECCV, 2025, pp. 20--37
2025
-
[94]
H. Chen, X. Wang, H. Chen, Z. Song, J. Jia, and W. Zhu, ``Grounding-prompter: Prompting llm with multimodal information for temporal sentence grounding in long videos,'' arXiv:2312.17117, 2023
2023 arXiv
-
[95]
Y. Xu, Y. Sun, Z. Xie, B. Zhai, and S. Du, ``Vtg-gpt: Tuning-free zero-shot video temporal grounding with gpt,'' Applied Sciences, vol. 14, no. 5, p. 1894, 2024
2024
-
[96]
H. Chen, X. Wang, H. Chen, Z. Zhang, W. Feng, B. Huang, J. Jia, and W. Zhu, ``Verified: A video corpus moment retrieval benchmark for fine-grained video understanding,'' arXiv:2410.08593, 2024
2024 arXiv
-
[97]
H. Lee, S. Hong, M. Sung, and J. Choi, ``Infusing environmental captions for long-form video language grounding,'' arXiv:2408.02336, 2024
2024 arXiv
-
[98]
D. Paul, M. R. Parvez, N. Mohammed, and S. Rahman, ``Videolights: Feature refinement and cross-task alignment transformer for joint video highlight detection and moment retrieval,'' arXiv:2412.01558, 2024
2024
-
[99]
Y. Sun, Y. Xu, Z. Xie, Y. Shu, and S. Du, ``Gptsee: Enhancing moment retrieval and highlight detection via description-based similarity features,'' SPL, vol. 31, pp. 521--525, 2024
2024
-
[100]
W. Liu, B. Miao, J. Cao, X. Zhu, B. Liu, M. Nasim, and A. Mian, ``Context-enhanced video moment retrieval with large language models,'' arXiv:2405.12540, 2024
2024 arXiv
-
[101]
P. Bao, C. Kong, Z. Shao, B. P. Ng, M. H. Er, and A. C. Kot, ``Vid-morp: Video moment retrieval pretraining from unlabeled videos in the wild,'' arXiv:2412.00811, 2024
2024 arXiv
-
[102]
Y. Xu, Y. Sun, B. Zhai, M. Li, W. Liang, Y. Li, and S. Du, ``Zero-shot video moment retrieval via off-the-shelf multimodal large language models,'' in AAAI, 2025, pp. 8978--8986
2025
-
[103]
S. Yu, J. Cho, P. Yadav, and M. Bansal, ``Self-chained image-language model for video localization and question answering,'' in NeurIPS, 2023, pp. 76\,749--76\,771
2023
-
[104]
K. Ma, X. Zang, Z. Feng, H. Fang, C. Ban, Y. Wei, Z. He, Y. Li, and H. Sun, ``Llavilo: Boosting video moment retrieval via adapter-based multimodal modeling,'' in ICCV, 2023, pp. 2798--2803
2023
-
[105]
Y. Liu, H. Hou, F. Ma, S. Ni, and F. R. Yu, ``Mllm-ta: Leveraging multimodal large language models for precise temporal video grounding,'' SPL, vol. 32, pp. 281--285, 2025
2025
-
[106]
W. Lu, J. Li, A. Yu, M.-C. Chang, S. Ji, and M. Xia, ``Llava-mr: Large language-and-vision assistant for video moment retrieval,'' arXiv:2411.14505, 2024
2024 arXiv
-
[107]
Y. Wu, X. Hu, Y. Sun, Y. Zhou, W. Zhu, F. Rao, B. Schiele, and X. Yang, ``Number it: Temporal grounding videos like flipping manga,'' arXiv:2411.10332, 2024
2024 arXiv
-
[108]
A. Deng, Z. Gao, A. Choudhuri, B. Planche, M. Zheng, B. Wang, T. Chen, C. Chen, and Z. Wu, ``Seq2time: Sequential knowledge transfer for video llm temporal grounding,'' arXiv:2411.16932, 2024
2024 arXiv
-
[109]
Y. Guo, J. Liu, M. Li, D. Cheng, X. Tang, D. Sui, Q. Liu, X. Chen, and K. Zhao, ``Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,'' in AAAI, 2025, pp. 3302--3310
2025
-
[110]
Z. Li, Q. Xu, D. Zhang, H. Song, Y. Cai, Q. Qi, R. Zhou, J. Pan, Z. Li, V. Tu et al., ``Groundinggpt: Language enhanced multi-modal grounding model,'' in ACL, 2024, pp. 6657--6678
2024
-
[111]
Hannan, M
T. Hannan, M. M. Islam, J. Gu, T. Seidl, and G. Bertasius, ``Revisionllm: Recursive vision-language model for temporal grounding in hour-long videos,'' arXiv:2411.14901, 2024
2024 arXiv
-
[112]
Huang, S
D.-A. Huang, S. Liao, S. Radhakrishnan, H. Yin, P. Molchanov, Z. Yu, and J. Kautz, ``Lita: Language instructed temporal-localization assistant,'' in ECCV, 2025, pp. 202--218
2025
-
[113]
X. Wang, F. Cheng, Z. Wang, H. Wang, M. M. Islam, L. Torresani, M. Bansal, G. Bertasius, and D. Crandall, ``Timerefine: Temporal grounding with time refining video llm,'' arXiv:2412.09601, 2024
2024 arXiv
-
[114]
H. Li, J. Chen, Z. Wei, S. Huang, T. Hui, J. Gao, X. Wei, and S. Liu, ``Llava-st: A multimodal large language model for fine-grained spatial-temporal understanding,'' arXiv:2501.08282, 2025
2025 arXiv
-
[115]
S. Chen, X. Lan, Y. Yuan, Z. Jie, and L. Ma, ``Timemarker: A versatile video-llm for long and short video understanding with superior temporal localization ability,'' arXiv:2411.18211, 2024
2024 arXiv
-
[116]
Q. Chen, S. Di, and W. Xie, ``Grounded multi-hop videoqa in long-form egocentric videos,'' in AAAI, 2025, pp. 2159--2167
2025
-
[117]
M. Nie, D. Ding, C. Wang, Y. Guo, J. Han, H. Xu, and L. Zhang, ``Slowfocus: Enhancing fine-grained temporal understanding in video llm,'' in NeurIPS, 2024
2024
-
[118]
Meinardus, A
B. Meinardus, A. Batra, A. Rohrbach, and M. Rohrbach, ``The surprising effectiveness of multimodal large language models for video moment retrieval,'' arXiv:2406.18113, 2024
2024
-
[119]
F. J. Fateh, U. Ahmed, H. Khan, M. Z. Zia, and Q.-H. Tran, ``Video llms for temporal reasoning in long videos,'' arXiv:2412.02930, 2024
2024 arXiv
-
[120]
Y. Wang, Y. Wang, P. Wu, J. Liang, D. Zhao, Y. Liu, and Z. Zheng, ``Efficient temporal extrapolation of multimodal large language models with temporal grounding bridge,'' in EMNLP, 2024, pp. 9972--9987
2024
-
[121]
X. Li, B. Wang, G. Shi, C. Feng, and J. Teng, ``Mitigating the discrepancy between video and text temporal sequences: A time-perception enhanced video grounding method for llm,'' in COLING, 2025, pp. 9804--9813
2025
-
[122]
Z. Yan, Z. Li, Y. He, C. Wang, K. Li, X. Li, X. Zeng, Z. Wang, Y. Wang, Y. Qiao et al., ``Task preference optimization: Improving multimodal large language models with vision task alignment,'' arXiv:2412.19326, 2024
2024 arXiv
-
[123]
J. Wang, Z. Liu, Y. Li, J. Ge, H. Xie, Y. Zhang et al., ``Spacevllm: Endowing multimodal large language model with spatio-temporal video grounding capability,'' arXiv:2503.13983, 2025
2025 arXiv
-
[124]
Z. Pang, M. Otani, and Y. Nakashima, ``Measure twice, cut once: Grasping video structures and event semantics with llms for video temporal localization,'' arXiv:2503.09027, 2025
2025
-
[125]
Y. Wang, Z. Wang, B. Xu, Y. Du, K. Lin, Z. Xiao, Z. Yue, J. Ju, L. Zhang, D. Yang, X. Fang, Z. He, Z. Luo, W. Wang, J. Lin, J. Luan, and Q. Jin, ``Time-r1: Post-training large vision language model for temporal video grounding,'' arXiv:2503.13377, 2025
2025 arXiv
-
[126]
Zhao, G.-P
H. Zhao, G.-P. Ji, R. Yan, H. Xiong, and Z. Li, ``Videoexpert: Augmented llm for temporal-sensitive video understanding,'' arXiv:2504.07519, 2025
2025 arXiv
-
[127]
X. Li, Z. Yan, D. Meng, L. Dong, X. Zeng, Y. He, Y. Wang, Y. Qiao, Y. Wang, and L. Wang, ``Videochat-r1: Enhancing spatio-temporal perception via reinforcement fine-tuning,'' arXiv:2504.06958, 2025
2025 arXiv
-
[128]
F. Luo, S. Lou, C. Chen, Z. Wang, C. Li, W. Shen, J. Guo, P. Li, M. Yan, J. Zhang et al., ``Museg: Reinforcing video temporal understanding via timestamp-aware multi-segment grounding,'' arXiv:2505.20715, 2025
2025 arXiv
-
[129]
Touvron, T
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi \`e re, N. Goyal, E. Hambro, F. Azhar et al., ``Llama: Open and efficient foundation language models,'' arXiv:2302.13971, 2023
2023 arXiv
-
[130]
Grauman, A
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu et al., ``Ego4d: Around the world in 3,000 hours of egocentric video,'' in CVPR, 2022, pp. 18\,995--19\,012
2022
-
[131]
[Online]
OpenAI, ``Gpt-4o,'' 2024. [Online]. Available: https://openai.com/index/hello-gpt-4o/
2024
-
[132]
G. Team, P. Georgiev, V. I. Lei, R. Burnell, L. Bai, A. Gulati, G. Tanzer, D. Vincent, Z. Pan, S. Wang et al., ``Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context,'' arXiv:2403.05530, 2024
2024 arXiv
-
[133]
M. Maaz, H. Rasheed, S. Khan, and F. S. Khan, ``Video-chatgpt: Towards detailed video understanding via large vision and language models,'' in ACL, 2024
2024
-
[134]
Reimers and I
N. Reimers and I. Gurevych, ``Sentence-bert: Sentence embeddings using siamese bert-networks,'' in EMNLP, 2019, pp. 3982--3992
2019
-
[135]
J. Chen, D. Zhu, X. Shen, X. Li, Z. Liu, P. Zhang, R. Krishnamoorthi, V. Chandra, Y. Xiong, and M. Elhoseiny, ``Minigpt-v2: large language model as a unified interface for vision-language multi-task learning,'' arXiv:2310.09478, 2023
2023 arXiv
-
[136]
A. Yang, B. Xiao, B. Wang, B. Zhang, C. Bian, C. Yin, C. Lv, D. Pan, D. Wang, D. Yan et al., ``Baichuan 2: Open large-scale language models,'' arXiv:2309.10305, 2023
2023 arXiv
-
[137]
K. C. Fraser and S. Kiritchenko, ``Examining gender and racial bias in large vision-language models using a novel dataset of parallel images,'' in EACL, 2024, pp. 690--713
2024
-
[138]
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu et al., ``Deepseekmath: Pushing the limits of mathematical reasoning in open language models,'' arXiv:2402.03300, 2024
2024 arXiv
-
[139]
Zellers, X
R. Zellers, X. Lu, J. Hessel, Y. Yu, J. S. Park, J. Cao, A. Farhadi, and Y. Choi, ``Merlot: Multimodal neural script knowledge models,'' in NeurIPS, 2021, pp. 23\,634--23\,651
2021
-
[140]
Y. Wang, Y. He, Y. Li, K. Li, J. Yu, X. Ma, X. Li, G. Chen, X. Chen, Y. Wang et al., ``Internvid: A large-scale video-text dataset for multimodal understanding and generation,'' arXiv:2307.06942, 2023
2023 arXiv
-
[141]
Zellers, J
R. Zellers, J. Lu, X. Lu, Y. Yu, Y. Zhao, M. Salehi, A. Kusupati, J. Hessel, A. Farhadi, and Y. Choi, ``Merlot reserve: Neural script knowledge through vision and language and sound,'' in CVPR, 2022, pp. 16\,375--16\,387
2022
-
[142]
Y. Zhai, S. Tong, X. Li, M. Cai, Q. Qu, Y. J. Lee, and Y. Ma, ``Investigating the catastrophic forgetting in multimodal large language models,'' arXiv:2309.10313, 2023
2023 arXiv
-
[143]
S. Chen, W. Jiang, W. Liu, and Y.-G. Jiang, ``Learning modality interaction for temporal sentence localization and event captioning in videos,'' in ECCV, 2020, pp. 333--351
2020
-
[144]
Wang, Z.-J
H. Wang, Z.-J. Zha, L. Li, D. Liu, and J. Luo, ``Structured multi-level interaction network for video moment localization via language query,'' in CVPR, 2021, pp. 7026--7035
2021
-
[145]
J. Wang, L. Ma, and W. Jiang, ``Temporally grounding language queries in videos by contextual boundary-aware prediction,'' in AAAI, 2020, pp. 12\,168--12\,175
2020
-
[146]
Chen, Y.-H
Y.-W. Chen, Y.-H. Tsai, and M.-H. Yang, ``End-to-end multi-modal video temporal grounding,'' in NeurIPS, 2021, pp. 28\,442--28\,453
2021
-
[147]
S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, and F. Wei, ``Beats: Audio pre-training with acoustic tokenizers,'' arXiv:2212.09058, 2022
2022 arXiv
-
[148]
Devlin, M.-W
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, ``Bert: Pre-training of deep bidirectional transformers for language understanding,'' in NAACL, 2019, pp. 4171--4186
2019
-
[149]
J. Gao, C. Sun, Z. Yang, and R. Nevatia, ``Tall: Temporal activity localization via language query,'' in ICCV, 2017, pp. 5267--5275
2017
-
[150]
L. Zhou, C. Xu, and J. Corso, ``Towards automatic learning of procedures from web instructional videos,'' in AAAI, 2018, pp. 7590--7598
2018
-
[151]
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta, ``Hollywood in homes: Crowdsourcing data collection for activity understanding,'' in ECCV, 2016, pp. 510--526
2016
-
[152]
Caba Heilbron, V
F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles, ``Activitynet: A large-scale video benchmark for human activity understanding,'' in CVPR, 2015, pp. 961--970
2015
-
[153]
J. Xiao, X. Shang, A. Yao, and T.-S. Chua, ``Next-qa: Next phase of question-answering to explaining temporal actions,'' in CVPR, 2021, pp. 9777--9786
2021
-
[154]
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell, ``Natural language object retrieval,'' in CVPR, 2016, pp. 4555--4564
2016
-
[155]
Fujita, T
S. Fujita, T. Hirao, H. Kamigaito, M. Okumura, and M. Nagata, ``Soda: Story oriented dense video captioning evaluation framework,'' in ECCV, 2020, pp. 517--531
2020
-
[156]
Banerjee and A
S. Banerjee and A. Lavie, ``Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,'' in ACL Workshop, 2005, pp. 65--72
2005
-
[157]
Vedantam, C
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, ``Cider: Consensus-based image description evaluation,'' in CVPR, 2015, pp. 4566--4575
2015
-
[158]
W. Liu, T. Mei, Y. Zhang, C. Che, and J. Luo, ``Multi-task deep visual-semantic embedding for video thumbnail selection,'' in CVPR, 2015, pp. 3707--3715
2015
-
[159]
Zhong, W
Y. Zhong, W. Ji, J. Xiao, Y. Li, W. Deng, and T.-S. Chua, ``Video question answering: Datasets, algorithms and challenges,'' in EMNLP, 2022, pp. 6439--6455
2022
-
[160]
X. Wang, Q. Si, J. Wu, S. Zhu, L. Cao, and L. Nie, ``Retake: Reducing temporal and knowledge redundancy for long video understanding,'' arXiv:2412.20504, 2024
2024 arXiv
-
[161]
Y. Li, C. Wang, and J. Jia, ``Llama-vid: An image is worth 2 tokens in large language models,'' in ECCV, 2024, pp. 323--340
2024
-
[162]
Viertola, V
I. Viertola, V. Iashin, and E. Rahtu, ``Temporally aligned audio for video with autoregression,'' in ICASSP, 2025, pp. 1--5
2025
-
[163]
Yariv, I
G. Yariv, I. Gat, S. Benaim, L. Wolf, I. Schwartz, and Y. Adi, ``Diverse and aligned audio-to-video generation via text-to-video model adaptation,'' in AAAI, 2024, pp. 6639--6647
2024
-
[164]
Radford, J
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, ``Robust speech recognition via large-scale weak supervision,'' in ICML, 2023, pp. 28\,492--28\,518
2023
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.