Pith. sign in

REVIEW 4 major objections 3 minor 3 cited by

A Survey on Video Temporal Grounding with Multimodal Large Language Model

T0 review · 4 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This survey argues that video temporal grounding methods built on multimodal large language models are surpassing traditional fine-tuned approaches, and organizes the field with a three-dimensional taxonomy.

desk verdict Useful-looking survey of VTG-MLLMs; only the abstract is legible, so the core empirical claim is unsupported in the artifact we have. read the letter →

arxiv 2508.10922 v1 pith:DAUTZFKX submitted 2025-08-07 cs.CV

classification cs.CV
keywords videotemporalgroundingmultimodallargelanguagemodelsurveytaxonomyzero-shotgeneralizationunderstandingbenchmarkmulti-tasklearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a systematic survey of video temporal grounding (VTG) methods built on multimodal large language models (MLLMs). It argues that these VTG-MLLM approaches are surpassing traditional fine-tuned methods, achieving competitive performance on standard benchmarks while generalizing better in zero-shot, multi-task, and multi-domain settings. To organize the field, the authors propose a three-dimensional taxonomy covering the functional role of the MLLM, the training paradigm, and video feature processing. The survey matters because a dedicated, comprehensive review of VTG-MLLMs was missing, and the taxonomy provides a shared map for positioning new work and identifying open problems.

What carries the argument

The key machinery is the three-dimensional taxonomy itself: (1) functional roles of MLLMs, highlighting architectural significance; (2) training paradigms for temporal reasoning and task adaptation; and (3) video feature processing techniques that determine spatiotemporal representation. This taxonomy is what allows the survey to compare methods systematically, summarize empirical findings, and point to limitations and research directions.

What would settle it

A concrete finding that would undermine the central claim: a published VTG-MLLM method that cannot be placed in any of the three taxonomy axes, or a rigorous multi-domain benchmark where traditional fine-tuned methods consistently beat MLLM-based methods.

Watch

Extended reading notes

Core claim

The central claim of the survey is that MLLM-based methods have become the leading paradigm for video temporal grounding. In the paper's own terms, VTG-MLLMs 'are gradually surpassing traditional fine-tuned methods': they stay competitive on standard benchmarks and excel in zero-shot, multi-task, and multi-domain generalization. The contribution is a three-dimensional taxonomy that classifies each approach by the role the MLLM plays, the training strategy used, and the video feature processing pipeline, together with a review of benchmarks, evaluation protocols, and empirical findings.

Load-bearing premise

The load-bearing premise is that the three-axis taxonomy cleanly classifies every VTG-MLLM approach and that the reported benchmark comparisons are representative of the field.

Editorial extensions

If this is right

  • Future VTG-MLLM papers can be positioned within a shared three-dimensional space, making method comparisons and reproducibility clearer.
  • The empirical trend implies that new VTG research should build on MLLMs rather than fine-tuned baselines, especially when targeting zero-shot or multi-domain deployment.
  • Benchmark design should include zero-shot, multi-task, and multi-domain settings to capture the claimed generalization advantage.
  • The survey's identified limitations and proposed directions provide a concrete agenda for improving temporal reasoning and task adaptation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The three-axis taxonomy may extend to neighboring video-language tasks such as video captioning or spatio-temporal grounding, since the axes are not specific to temporal localization.
  • The claim that MLLM methods surpass fine-tuned methods may depend on benchmark and domain choice; a standardized, broader evaluation would test how universal the advantage is.
  • In practice the three axes may overlap (training paradigms can influence feature processing), so the taxonomy is best read as complementary lenses rather than a strict partition.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. This manuscript is a survey of video temporal grounding with multimodal large language models (VTG-MLLMs). The abstract claims that VTG-MLLMs are gradually surpassing traditional fine-tuned methods, both in competitive performance and in generalization across zero-shot, multi-task, and multi-domain settings. The survey is organized around a three-dimensional taxonomy: functional roles of MLLMs, training paradigms, and video feature processing techniques. It also promises to cover benchmark datasets, evaluation protocols, empirical findings, limitations, and future research directions, with an accompanying GitHub repository. The supplied full text, however, is almost entirely composed of Unicode replacement characters, making the body and tables unreadable. Consequently, the survey's substantive content—its taxonomies, classifications, benchmark comparisons, and critical discussion—cannot be assessed from the provided artifact; only the abstract and the presence of an unreadable appendix/table section are available.

Significance. If the claims in the abstract were fully supported by the body, the survey would make a useful contribution by providing an organizing three-axis taxonomy for a rapidly growing but scattered literature, and by synthesizing empirical evidence on whether MLLM-based approaches truly outperform traditional fine-tuned methods. However, the significance of the empirical claim depends entirely on the completeness, accuracy, and protocol-matching of benchmark comparisons, which are not visible in the supplied full text. The taxonomy's value similarly depends on whether the three axes are disjoint and jointly exhaustive, a property that is not established in the abstract and cannot be checked. No machine-checked proofs, reproducible code, or parameter-free derivations are present. The manuscript's potential significance is therefore real but currently unverifiable from the artifact provided.

major comments (4)
  1. [Abstract vs. Full Text] The central empirical claim, that 'VTG approaches based on MLLMs are gradually surpassing traditional fine-tuned methods' and generalize better across zero-shot, multi-task, and multi-domain settings, is load-bearing. In the supplied full text, the body after the abstract is overwhelmingly composed of U+FFFD replacement characters, and the tables on the later pages are unreadable. No benchmark numbers, evaluation protocols, error bars, or controlled comparisons are visible. This is a verification gap rather than an internal inconsistency, but it means the survey's main empirical conclusion is unsupported in the provided artifact. Please provide a readable manuscript with the actual text, tables, and references so the evidence can be assessed.
  2. [Three-dimensional taxonomy] The abstract introduces a three-axis taxonomy—functional roles of MLLMs, training paradigms, and video feature processing—as the organizing structure. The taxonomy is only stated, not demonstrated to be exhaustive or disjoint. In particular, 'functional role' and 'training paradigm' may overlap: a method's functional role is often expressed through its training objective or adapter design. Without explicit definitions of category boundaries and a classification of representative methods, a reader cannot judge whether the taxonomy systematically covers all current VTG-MLLM work or whether some methods fall between categories. This is a load-bearing organizational claim that requires concrete support in the survey body.
  3. [Benchmarks section] The abstract promises discussion of benchmark datasets, evaluation protocols, and empirical findings. In the supplied full text, the only readable section appears to be a table-like block of replacement characters, which cannot be interpreted. There is no verifiable presentation of datasets, metrics, or results. Since the survey's contribution includes 'summarize empirical findings,' the absence of legible benchmark analysis prevents any validation of the claimed performance trends. Please provide the benchmark tables in machine-readable or otherwise legible form, with protocol details and citations.
  4. [Limitations and future directions] The abstract states that the survey 'identifies existing limitations and proposes promising research directions.' These sections are also unreadable in the supplied artifact. A survey's critical value depends on its assessment of limitations (e.g., evaluation biases, benchmark saturation, generalization gaps). Without access to these sections, the completeness and even-handedness of the survey cannot be judged.
minor comments (3)
  1. [Full text encoding] The supplied PDF/HTML appears to have a text-encoding failure: almost all body text is replaced by U+FFFD replacement characters. This must be fixed in the submission; a clean, readable version is essential for review and for reader use.
  2. [Repository link] The abstract refers to https://github.com/ki-lw/Awesome-MLLMs-for-Video-Temporal-Grounding for additional resources. The repository may be useful, but it is not a substitute for the survey's own text and tables. Please ensure the survey body stands alone independently of the repository.
  3. [References] The provided text does not show a reference list. A survey must include complete citations for all discussed methods and benchmarks; these need to be legible and correct.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity identified: this is a survey with no derivation chain; the main empirical claim is unverifiable from the garbled body but is not shown to reduce to its own inputs.

full rationale

The paper is a survey of video temporal grounding with multimodal large language models. Its contribution is a three-dimensional taxonomy (functional roles of MLLMs, training paradigms, video feature processing) and an empirical synthesis claiming that VTG-MLLM methods are 'gradually surpassing traditional fine-tuned methods' and generalize better in zero-shot, multi-task, and multi-domain settings. A survey does not ordinarily present a derivation chain of the kind that could be circular: there are no equations, no fitted parameters renamed as predictions, and no formal claims that reduce to their definitions. The abstract's empirical claim depends on benchmark evidence, and the supplied full text is largely unreadable U+FFFD replacement characters, so those benchmark tables cannot be audited. However, unverifiability is not circularity. The taxonomy's axes may be overlapping (a method's functional role is often expressed through its training paradigm), which is a classification-quality concern rather than a circular-reasoning concern. No load-bearing self-citation can be identified from the provided artifact. Under the hard rules requiring a quotable specific reduction, no circular step can be exhibited. The honest finding is therefore no significant circularity, score 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

As a survey, the paper introduces no free parameters and no invented entities. Its load-bearing assumptions are domain assumptions about taxonomy completeness and faithful reporting of benchmark comparisons, both asserted in the abstract and unverifiable from the abstract alone.

assumptions (2)
  • domain assumption The three-axis taxonomy (functional roles, training paradigms, video feature processing) is exhaustive and categories are disjoint.
    The abstract asserts a three-dimensional taxonomy, but completeness and mutual exclusivity cannot be verified without the full text.
  • domain assumption The empirical claim that VTG-MLLMs surpass traditional fine-tuned methods is faithfully reported from the cited benchmarks.
    Stated in the abstract; with only the abstract available, the underlying numbers and comparison configurations cannot be checked.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey on Video Temporal Grounding with Multimodal Large Language Model." pith.science (2026). https://pith.science/paper/DAUTZFKX

@misc{pith2026250810922,
  author       = {Pith},
  title        = {Pith review of: A Survey on Video Temporal Grounding with Multimodal Large Language Model},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DAUTZFKX}},
  note         = {Machine review of arXiv:2508.10922}
}
read the original abstract

The recent advancement in video temporal grounding (VTG) has significantly enhanced fine-grained video understanding, primarily driven by multimodal large language models (MLLMs). With superior multimodal comprehension and reasoning abilities, VTG approaches based on MLLMs (VTG-MLLMs) are gradually surpassing traditional fine-tuned methods. They not only achieve competitive performance but also excel in generalization across zero-shot, multi-task, and multi-domain settings. Despite extensive surveys on general video-language understanding, comprehensive reviews specifically addressing VTG-MLLMs remain scarce. To fill this gap, this survey systematically examines current research on VTG-MLLMs through a three-dimensional taxonomy: 1) the functional roles of MLLMs, highlighting their architectural significance; 2) training paradigms, analyzing strategies for temporal reasoning and task adaptation; and 3) video feature processing techniques, which determine spatiotemporal representation effectiveness. We further discuss benchmark datasets, evaluation protocols, and summarize empirical findings. Finally, we identify existing limitations and propose promising research directions. For additional resources and details, readers are encouraged to visit our repository at https://github.com/ki-lw/Awesome-MLLMs-for-Video-Temporal-Grounding.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. NEST: Narrative Event Structures in Time for Long Video Understanding

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    NEST is a new benchmark dataset for narrative event structures in long videos, with baselines reporting ETD below 8%, EL under 6%, EAE below 11%, and ERE at 35-44% F1.

  2. Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition

    cs.CV 2026-06 conditional novelty 7.0 of 10

    Hour-long video temporal grounding is a search problem, shown by a new benchmark where all Video-LLMs collapse, frame retrieval outperforms them, 85% of failures are search-related, and a retrieve-then-ground hybrid i...

  3. Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs

    cs.CV 2025-10 conditional novelty 6.0 of 10

    Selecting the visual input that minimizes an MLLM's output entropy (or maximizes its yes/no confidence) improves fine-grained visual search, long-video QA, and temporal grounding without any training.

Reference graph

Works this paper leans on

164 extracted references · 34 canonical work pages · cited by 3 Pith papers

  1. [1]

    11em plus .33em minus .07em 4000 4000 100 4000 4000 500 `\.=1000 = #1 \@IEEEnotcompsoconly \@IEEEcompsoconly #1 * [1] 0pt [0pt][0pt] #1 * [1] 0pt [0pt][0pt] #1 * \| ** #1 \@IEEEauthorblockNstyle \@IEEEcompsocnotconfonly \@IEEEauthorblockAstyle \@IEEEcompsocnotconfonly \@IEEEcompsocconfonly \@IEEEauthordefaulttextstyle \@IEEEcompsocnotconfonly \@IEEEauthor...

  2. [2]

    Y. Wang, K. Li, Y. Li, Y. He, B. Huang, Z. Zhao, H. Zhang, J. Xu, Y. Liu, Z. Wang et al., ``Internvideo: General video foundation models via generative and discriminative learning,'' arXiv:2212.03191, 2022

  3. [3]

    Y. Wang, K. Li, X. Li, J. Yu, Y. He, G. Chen, B. Pei, R. Zheng, Z. Wang, Y. Shi et al., ``Internvideo2: Scaling foundation models for multimodal video understanding,'' in ECCV, 2024, pp. 396--416

  4. [4]

    M. Wang, J. Xing, J. Mei, Y. Liu, and Y. Jiang, ``Actionclip: Adapting language-image pretrained models for video action recognition,'' TNNLS, vol. 36, no. 1, pp. 625--637, 2023

  5. [5]

    A. Yang, A. Nagrani, P. H. Seo, A. Miech, J. Pont-Tuset, I. Laptev, J. Sivic, and C. Schmid, ``Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,'' in CVPR, 2023, pp. 10\,714--10\,726

  6. [6]

    Regneri, M

    M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal, ``Grounding action descriptions in videos,'' TACL, vol. 1, pp. 25--36, 2013

  7. [7]

    Anne Hendricks, O

    L. Anne Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell, ``Localizing moments in video with natural language,'' in ICCV, 2017, pp. 5803--5812

  8. [8]

    Krishna, K

    R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. Carlos Niebles, ``Dense-captioning events in videos,'' in ICCV, 2017, pp. 706--715

Show all 164 references
  1. [9]

    T. Wang, R. Zhang, Z. Lu, F. Zheng, R. Cheng, and P. Luo, ``End-to-end dense video captioning with parallel decoding,'' in ICCV, 2021, pp. 6847--6857

  2. [10]

    Y. Song, J. Vallmitjana, A. Stent, and A. Jaimes, ``Tvsum: Summarizing web videos using titles,'' in CVPR, 2015, pp. 5179--5187

  3. [11]

    Xiong, Y

    B. Xiong, Y. Kalantidis, D. Ghadiyaram, and K. Grauman, ``Less is more: Learning highlight detection from video duration,'' in CVPR, 2019, pp. 1258--1267

  4. [12]

    J. Xiao, A. Yao, Y. Li, and T.-S. Chua, ``Can i trust your answer? visually grounded video question answering,'' in CVPR, 2024, pp. 13\,204--13\,214

  5. [13]

    Chen, Y.-C

    J.-J. Chen, Y.-C. Liao, H.-C. Lin, Y.-C. Yu, Y.-C. Chen, and F. Wang, ``Rextime: A benchmark suite for reasoning-across-time in videos,'' in NeurIPS, 2024, pp. 28\,662--28\,673

  6. [14]

    Zhang, X

    D. Zhang, X. Dai, X. Wang, Y.-F. Wang, and L. S. Davis, ``Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment,'' in CVPR, 2019, pp. 1247--1257

  7. [15]

    W. Moon, S. Hyun, S. Park, D. Park, and J.-P. Heo, ``Query-dependent video representation for moment retrieval and highlight detection,'' in CVPR, 2023, pp. 23\,023--23\,033

  8. [16]

    M. Ma, S. Yoon, J. Kim, Y. Lee, S. Kang, and C. D. Yoo, ``Vlanet: Video-language alignment network for weakly-supervised video moment retrieval,'' in ECCV, 2020, pp. 156--171

  9. [17]

    Zhang, Y

    M. Zhang, Y. Yang, X. Chen, Y. Ji, X. Xu, J. Li, and H. T. Shen, ``Multi-stage aggregated transformer network for temporal language localization in videos,'' in CVPR, 2021, pp. 12\,669--12\,678

  10. [18]

    Y. Liu, S. Li, Y. Wu, C.-W. Chen, Y. Shan, and X. Qie, ``Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection,'' in CVPR, 2022, pp. 3042--3051

  11. [19]

    Y. Yuan, T. Mei, and W. Zhu, ``To find where you talk: Temporal sentence localization in video with attention based location regression,'' in AAAI, 2019, pp. 9159--9166

  12. [20]

    Zhang, S

    S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin et al., ``Opt: Open pre-trained transformer language models,'' arXiv:2205.01068, 2022

  13. [21]

    H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma et al., ``Scaling instruction-finetuned language models,'' JMLR, vol. 25, no. 70, pp. 1--53, 2024

  14. [22]

    D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi et al., ``Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,'' arXiv:2501.12948, 2025

  15. [23]

    H. Liu, C. Li, Q. Wu, and Y. J. Lee, ``Visual instruction tuning,'' in NeurIPS, 2023, pp. 34\,892--34\,916

  16. [24]

    W. Dai, J. Li, D. LI, A. Tiong, J. Zhao, W. Wang, B. Li, P. N. Fung, and S. Hoi, ``Instructblip: Towards general-purpose vision-language models with instruction tuning,'' in NeurIPS, 2023, pp. 49\,250--49\,267

  17. [25]

    Cheng, S

    Z. Cheng, S. Leng, H. Zhang, Y. Xin, X. Li, G. Chen, Y. Zhu, W. Zhang, Z. Luo, D. Zhao, and L. Bing, ``Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms,'' arXiv:2406.07476, 2024

  18. [26]

    L. Xu, Y. Zhao, D. Zhou, Z. Lin, S. K. Ng, and J. Feng, ``Pllava: Parameter-free llava extension from images to videos for video dense captioning,'' arXiv:2404.16994, 2024

  19. [27]

    Huang, X

    B. Huang, X. Wang, H. Chen, Z. Song, and W. Zhu, ``Vtimellm: Empower llm to grasp video moments,'' in CVPR, 2024, pp. 14\,271--14\,280

  20. [28]

    S. Ren, L. Yao, S. Li, X. Sun, and L. Hou, ``Timechat: A time-sensitive multimodal large language model for long video understanding,'' in CVPR, 2024, pp. 14\,313--14\,323

  21. [29]

    L. Qian, J. Li, Y. Wu, Y. Ye, H. Fei, T.-S. Chua, Y. Zhuang, and S. Tang, ``Momentor: Advancing video large language model with fine-grained temporal reasoning,'' in ICML, 2024, pp. 41\,340--41\,356

  22. [30]

    M. Qu, X. Chen, W. Liu, A. Li, and Y. Zhao, ``Chatvtg: Video temporal grounding via chat with video dialogue large language models,'' in CVPR, 2024, pp. 1847--1856

  23. [31]

    H. Qin, J. Xiao, and A. Yao, ``Question-answering dense video events,'' arXiv:2409.04388, 2024

  24. [32]

    Y. Guo, J. Liu, M. Li, X. Tang, Q. Liu, and X. Chen, ``Trace: Temporal grounding video llm via causal event modeling,'' in ICLR, 2025

  25. [33]

    Y. Wang, X. Meng, J. Liang, Y. Wang, Q. Liu, and D. Zhao, ``Hawkeye: Training video-text llms for grounding text in videos,'' arXiv:2403.10228, 2024

  26. [34]

    X. Zeng, K. Li, C. Wang, X. Li, T. Jiang, Z. Yan, S. Li, Y. Shi, Z. Yue, Y. Wang et al., ``Timesuite: Improving mllms for long video understanding via grounded tuning,'' in ICLR, 2025

  27. [35]

    H. Wang, Z. Xu, Y. Cheng, S. Diao, Y. Zhou, Y. Cao, Q. Wang, W. Ge, and L. Huang, ``Grounded-videollm: Sharpening fine-grained temporal grounding in video large language models,'' arXiv:2410.03290, 2024

  28. [36]

    Y. Liu, K. Q. Lin, C. W. Chen, and M. Z. Shou, ``Videomind: A chain-of-lora agent for long video reasoning,'' arXiv:2503.13444, 2025

  29. [37]

    C. Li, Z. Gan, Z. Yang, J. Yang, L. Li, L. Wang, J. Gao et al., ``Multimodal foundation models: From specialists to general-purpose assistants,'' arXiv:2309.10020, 2023

  30. [38]

    Y. Zhu, X. Li, C. Liu, M. Zolfaghari, Y. Xiong, C. Wu, Z. Zhang, J. Tighe, R. Manmatha, and M. Li, ``A comprehensive study of deep video action recognition,'' arXiv:2012.06567, 2020

  31. [39]

    Abdar, M

    M. Abdar, M. Kollati, S. Kuraparthi, F. Pourpanah, D. McDuff, M. Ghavamzadeh, S. Yan, A. Mohamed, A. Khosravi, E. Cambria et al., ``A review of deep learning for video captioning,'' TPAMI, pp. 1--20, 2024

  32. [40]

    Z. Xing, Q. Feng, H. Chen, Q. Dai, H. Hu, H. Xu, Z. Wu, and Y.-G. Jiang, ``A survey on video diffusion models,'' ACM Computing Surveys, vol. 57, no. 2, pp. 1--42, 2024

  33. [41]

    Zhang, J

    J. Zhang, J. Huang, S. Jin, and S. Lu, ``Vision-language models for vision tasks: A survey,'' TPAMI, vol. 46, no. 8, pp. 5625--5644, 2024

  34. [42]

    X. Liu, X. Nie, Z. Tan, J. Guo, and Y. Yin, ``A survey on natural language video localization,'' arXiv:2104.00234, 2021

  35. [43]

    Y. Yang, Z. Li, and G. Zeng, ``A survey of temporal activity localization via language in untrimmed videos,'' in ICCST, 2020, pp. 596--601

  36. [44]

    X. Lan, Y. Yuan, X. Wang, Z. Wang, and W. Zhu, ``A survey on temporal sentence grounding in videos,'' TOMCCAP, vol. 19, no. 2, pp. 1--33, 2023

  37. [45]

    Zhang, A

    H. Zhang, A. Sun, W. Jing, and J. T. Zhou, ``Temporal sentence grounding in videos: A survey and future directions,'' TPAMI, vol. 45, no. 8, pp. 10\,443--10\,465, 2023

  38. [46]

    Zhang, H

    S. Zhang, H. Peng, J. Fu, and J. Luo, ``Learning 2d temporal adjacent networks for moment localization with natural language,'' in AAAI, 2020, pp. 12\,870--12\,877

  39. [47]

    K. Q. Lin, P. Zhang, J. Chen, S. Pramanick, D. Gao, A. J. Wang, R. Yan, and M. Z. Shou, ``Univtg: Towards unified video-language temporal grounding,'' in ICCV, 2023, pp. 2794--2804

  40. [48]

    W. Yang, T. Zhang, Y. Zhang, and F. Wu, ``Local correspondence network for weakly supervised temporal sentence grounding,'' TIP, vol. 30, pp. 3252--3262, 2021

  41. [49]

    D. Liu, X. Qu, J. Dong, P. Zhou, Y. Cheng, W. Wei, Z. Xu, and Y. Xie, ``Context-aware biaffine localizing network for temporal sentence grounding,'' in CVPR, 2021, pp. 11\,235--11\,244

  42. [50]

    S. Xiao, L. Chen, S. Zhang, W. Ji, J. Shao, L. Ye, and J. Xiao, ``Boundary proposal network for two-stage natural language video localization,'' in AAAI, 2021, pp. 2986--2994

  43. [51]

    Zhang, A

    H. Zhang, A. Sun, W. Jing, L. Zhen, J. T. Zhou, and R. S. M. Goh, ``Natural language video localization: A revisit in span-based question answering framework,'' TPAMI, vol. 44, no. 8, pp. 4252--4266, 2021

  44. [52]

    Y. Liu, Z. Ma, Z. Qi, Y. Wu, Y. Shan, and C. W. Chen, ``Et bench: Towards open-ended event-level video-language understanding,'' in NeurIPS, 2024, pp. 32\,076--32\,110

  45. [53]

    L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong, ``End-to-end dense video captioning with masked transformer,'' in CVPR, 2018, pp. 8739--8748

  46. [54]

    Iashin and E

    V. Iashin and E. Rahtu, ``Multi-modal dense video captioning,'' in CVPRW, 2020, pp. 958--959

  47. [55]

    J. Lei, T. L. Berg, and M. Bansal, ``Detecting moments and highlights in videos via natural language queries,'' in NeurIPS, 2021, pp. 11\,846--11\,858

  48. [56]

    G. Chen, Y. Liu, Y. Huang, Y. He, B. Pei, J. Xu, Y. Wang, T. Lu, and L. Wang, ``Cg-bench: Clue-grounded question answering benchmark for long video understanding,'' arXiv:2412.12075, 2024

  49. [57]

    H. Liu, X. Ma, C. Zhong, Y. Zhang, and W. Lin, ``Timecraft: Navigate weakly-supervised temporal grounded video question answering via bi-directional reasoning,'' in ECCV, 2025, pp. 92--107

  50. [58]

    Zhang, X

    H. Zhang, X. Li, and L. Bing, ``Video-llama: An instruction-tuned audio-visual language model for video understanding,'' in EMNLP, 2023, pp. 543--553

  51. [59]

    B. Lin, Y. Ye, B. Zhu, J. Cui, M. Ning, P. Jin, and L. Yuan, ``Video-llava: Learning united visual representation by alignment before projection,'' arXiv:2311.10122, 2023

  52. [60]

    K. Li, Y. He, Y. Wang, Y. Li, W. Wang, P. Luo, Y. Wang, L. Wang, and Y. Qiao, ``Videochat: Chat-centric video understanding,'' arXiv:2305.06355, 2023

  53. [61]

    Dosovitskiy, L

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., ``An image is worth 16x16 words: Transformers for image recognition at scale,'' in ICLR, 2020

  54. [62]

    Q. Sun, Y. Fang, L. Wu, X. Wang, and Y. Cao, ``Eva-clip: Improved training techniques for clip at scale,'' arXiv:2303.15389, 2023

  55. [63]

    Y. Fang, W. Wang, B. Xie, Q. Sun, L. Wu, X. Wang, T. Huang, X. Wang, and Y. Cao, ``Eva: Exploring the limits of masked visual representation learning at scale,'' in CVPR, 2023, pp. 19\,358--19\,369

  56. [64]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., ``Learning transferable visual models from natural language supervision,'' in ICML, 2021, pp. 8748--8763

  57. [65]

    Feichtenhofer, Y

    C. Feichtenhofer, Y. Li, K. He et al., ``Masked autoencoders as spatiotemporal learners,'' in NeurIPS, 2022, pp. 35\,946--35\,958

  58. [66]

    Z. Tong, Y. Song, J. Wang, and L. Wang, ``Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training,'' in NeurIPS, 2022, pp. 10\,078--10\,093

  59. [67]

    L. Wang, B. Huang, Z. Zhao, Z. Tong, Y. He, Y. Wang, Y. Wang, and Y. Qiao, ``Videomae v2: Scaling video masked autoencoders with dual masking,'' in CVPR, 2023, pp. 14\,549--14\,560

  60. [68]

    J. Wang, Y. Ge, R. Yan, Y. Ge, K. Q. Lin, S. Tsutsui, X. Lin, G. Cai, J. Wu, Y. Shan et al., ``All in one: Exploring unified video-language pre-training,'' in CVPR, 2023, pp. 6598--6608

  61. [69]

    G. Sun, W. Yu, C. Tang, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, Y. Wang, and C. Zhang, ``video-salmonn: Speech-enhanced audio-visual large language models,'' in ICML, 2024, pp. 47\,198--47\,217

  62. [70]

    S. Azad, V. Vineet, and Y. S. Rawat, ``Hierarq: Task-aware hierarchical q-former for enhanced video understanding,'' arXiv:2503.08585, 2025

  63. [71]

    Alayrac, J

    J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al., ``Flamingo: a visual language model for few-shot learning,'' in NeurIPS, 2022, pp. 23\,716--23\,736

  64. [72]

    J. Li, D. Li, S. Savarese, and S. Hoi, ``Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,'' in ICML, 2023, pp. 19\,730--19\,742

  65. [73]

    Q. Ye, H. Xu, G. Xu, J. Ye, M. Yan, Y. Zhou, J. Wang, A. Hu, P. Shi, Y. Shi et al., ``mplug-owl: Modularization empowers large language models with multimodality,'' arXiv:2304.14178, 2023

  66. [74]

    Huang, L

    S. Huang, L. Dong, W. Wang, Y. Hao, S. Singhal, S. Ma, T. Lv, L. Cui, O. K. Mohammed, B. Patra et al., ``Language is not all you need: Aligning perception with language models,'' in NeurIPS, 2023, pp. 72\,096--72\,109

  67. [75]

    H. Liu, C. Li, Y. Li, and Y. J. Lee, ``Improved baselines with visual instruction tuning,'' in CVPR, 2024, pp. 26\,296--26\,306

  68. [76]

    F. Li, R. Zhang, H. Zhang, Y. Zhang, B. Li, W. Li, Z. Ma, and C. Li, ``Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models,'' arXiv:2407.07895, 2024

  69. [77]

    J. Ye, H. Xu, H. Liu, A. Hu, M. Yan, Q. Qian, J. Zhang, F. Huang, and J. Zhou, ``mplug-owl3: Towards long image-sequence understanding in multi-modal large language models,'' arXiv:2408.04840, 2024

  70. [78]

    K. Chen, D. Shen, H. Zhong, H. Zhong, K. Xia, D. Xu, W. Yuan, Y. Hu, B. Wen, T. Zhang et al., ``Evlm: An efficient vision-language model for visual understanding,'' arXiv:2407.14177, 2024

  71. [79]

    M. Shi, S. Wang, C.-Y. Chen, J. Jain, K. Wang, J. Xiong, G. Liu, Z. Yu, and H. Shi, ``Slow-fast architecture for video multi-modal large language models,'' arXiv:2504.01328, 2025

  72. [80]

    Miech, D

    A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic, ``Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,'' in ICCV, 2019, pp. 2630--2640

  73. [81]

    M. Bain, A. Nagrani, G. Varol, and A. Zisserman, ``Frozen in time: A joint video and image encoder for end-to-end retrieval,'' in ICCV, 2021, pp. 1728--1738

  74. [82]

    Schuhmann, R

    C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman et al., ``Laion-5b: An open large-scale dataset for training next generation image-text models,'' in NeurIPS, 2022, pp. 25\,278--25\,294

  75. [83]

    L. Li, Y. Yin, S. Li, L. Chen, P. Wang, S. Ren, M. Li, Y. Yang, J. Xu, X. Sun et al., ``M ^3 it: A large-scale dataset towards multi-modal multilingual instruction tuning,'' arXiv:2306.04387, 2023

  76. [84]

    Z. Yin, J. Wang, J. Cao, Z. Shi, D. Liu, M. Li, X. Huang, Z. Wang, L. Sheng, L. Bai et al., ``Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark,'' in NeurIPS, 2023, pp. 26\,650--26\,685

  77. [85]

    H. Shao, S. Qian, H. Xiao, G. Song, Z. Zong, L. Wang, Y. Liu, and H. Li, ``Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning,'' in NeurIPS, 2024, pp. 8612--8642

  78. [86]

    T. Gao, P. Chen, M. Zhang, C. Fu, Y. Shen, Y. Zhang, S. Zhang, X. Zheng, X. Sun, L. Cao et al., ``Cantor: Inspiring multimodal chain-of-thought of mllm,'' in ACM MM, 2024, pp. 9096--9105

  79. [87]

    J. Wu, Z. Zhang, Y. Xia, X. Li, Z. Xia, A. Chang, T. Yu, S. Kim, R. A. Rossi, R. Zhang et al., ``Visual prompting in multimodal large language models: A survey,'' arXiv:2409.15310, 2024

  80. [88]

    E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen et al., ``Lora: Low-rank adaptation of large language models,'' in ICLR, 2022

  81. [89]

    R. Pan, X. Liu, S. Diao, R. Pi, J. Zhang, C. Han, and T. Zhang, ``Lisa: layerwise importance sampling for memory-efficient large language model fine-tuning,'' in NeurIPS, 2024, pp. 57\,018--57\,049

  82. [90]

    Liu, C.-Y

    S.-Y. Liu, C.-Y. Wang, H. Yin, P. Molchanov, Y.-C. F. Wang, K.-T. Cheng, and M.-H. Chen, ``Dora: Weight-decomposed low-rank adaptation,'' in ICML, 2024, pp. 32\,100--32\,121

  83. [91]

    W. Cai, J. Huang, S. Gong, H. Jin, and Y. Liu, ``Mllm as video narrator: Mitigating modality imbalance in video moment retrieval,'' arXiv:2406.17880, 2024

  84. [92]

    Di and W

    S. Di and W. Xie, ``Grounded question-answering in long egocentric videos,'' in CVPR, 2024, pp. 12\,934--12\,943

  85. [93]

    Zheng, X

    M. Zheng, X. Cai, Q. Chen, Y. Peng, and Y. Liu, ``Training-free video temporal grounding using large-scale pre-trained models,'' in ECCV, 2025, pp. 20--37

  86. [94]

    H. Chen, X. Wang, H. Chen, Z. Song, J. Jia, and W. Zhu, ``Grounding-prompter: Prompting llm with multimodal information for temporal sentence grounding in long videos,'' arXiv:2312.17117, 2023

  87. [95]

    Y. Xu, Y. Sun, Z. Xie, B. Zhai, and S. Du, ``Vtg-gpt: Tuning-free zero-shot video temporal grounding with gpt,'' Applied Sciences, vol. 14, no. 5, p. 1894, 2024

  88. [96]

    H. Chen, X. Wang, H. Chen, Z. Zhang, W. Feng, B. Huang, J. Jia, and W. Zhu, ``Verified: A video corpus moment retrieval benchmark for fine-grained video understanding,'' arXiv:2410.08593, 2024

  89. [97]

    H. Lee, S. Hong, M. Sung, and J. Choi, ``Infusing environmental captions for long-form video language grounding,'' arXiv:2408.02336, 2024

  90. [98]

    D. Paul, M. R. Parvez, N. Mohammed, and S. Rahman, ``Videolights: Feature refinement and cross-task alignment transformer for joint video highlight detection and moment retrieval,'' arXiv:2412.01558, 2024

  91. [99]

    Y. Sun, Y. Xu, Z. Xie, Y. Shu, and S. Du, ``Gptsee: Enhancing moment retrieval and highlight detection via description-based similarity features,'' SPL, vol. 31, pp. 521--525, 2024

  92. [100]

    W. Liu, B. Miao, J. Cao, X. Zhu, B. Liu, M. Nasim, and A. Mian, ``Context-enhanced video moment retrieval with large language models,'' arXiv:2405.12540, 2024

  93. [101]

    P. Bao, C. Kong, Z. Shao, B. P. Ng, M. H. Er, and A. C. Kot, ``Vid-morp: Video moment retrieval pretraining from unlabeled videos in the wild,'' arXiv:2412.00811, 2024

  94. [102]

    Y. Xu, Y. Sun, B. Zhai, M. Li, W. Liang, Y. Li, and S. Du, ``Zero-shot video moment retrieval via off-the-shelf multimodal large language models,'' in AAAI, 2025, pp. 8978--8986

  95. [103]

    S. Yu, J. Cho, P. Yadav, and M. Bansal, ``Self-chained image-language model for video localization and question answering,'' in NeurIPS, 2023, pp. 76\,749--76\,771

  96. [104]

    K. Ma, X. Zang, Z. Feng, H. Fang, C. Ban, Y. Wei, Z. He, Y. Li, and H. Sun, ``Llavilo: Boosting video moment retrieval via adapter-based multimodal modeling,'' in ICCV, 2023, pp. 2798--2803

  97. [105]

    Y. Liu, H. Hou, F. Ma, S. Ni, and F. R. Yu, ``Mllm-ta: Leveraging multimodal large language models for precise temporal video grounding,'' SPL, vol. 32, pp. 281--285, 2025

  98. [106]

    W. Lu, J. Li, A. Yu, M.-C. Chang, S. Ji, and M. Xia, ``Llava-mr: Large language-and-vision assistant for video moment retrieval,'' arXiv:2411.14505, 2024

  99. [107]

    Y. Wu, X. Hu, Y. Sun, Y. Zhou, W. Zhu, F. Rao, B. Schiele, and X. Yang, ``Number it: Temporal grounding videos like flipping manga,'' arXiv:2411.10332, 2024

  100. [108]

    A. Deng, Z. Gao, A. Choudhuri, B. Planche, M. Zheng, B. Wang, T. Chen, C. Chen, and Z. Wu, ``Seq2time: Sequential knowledge transfer for video llm temporal grounding,'' arXiv:2411.16932, 2024

  101. [109]

    Y. Guo, J. Liu, M. Li, D. Cheng, X. Tang, D. Sui, Q. Liu, X. Chen, and K. Zhao, ``Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,'' in AAAI, 2025, pp. 3302--3310

  102. [110]

    Z. Li, Q. Xu, D. Zhang, H. Song, Y. Cai, Q. Qi, R. Zhou, J. Pan, Z. Li, V. Tu et al., ``Groundinggpt: Language enhanced multi-modal grounding model,'' in ACL, 2024, pp. 6657--6678

  103. [111]

    Hannan, M

    T. Hannan, M. M. Islam, J. Gu, T. Seidl, and G. Bertasius, ``Revisionllm: Recursive vision-language model for temporal grounding in hour-long videos,'' arXiv:2411.14901, 2024

  104. [112]

    Huang, S

    D.-A. Huang, S. Liao, S. Radhakrishnan, H. Yin, P. Molchanov, Z. Yu, and J. Kautz, ``Lita: Language instructed temporal-localization assistant,'' in ECCV, 2025, pp. 202--218

  105. [113]

    X. Wang, F. Cheng, Z. Wang, H. Wang, M. M. Islam, L. Torresani, M. Bansal, G. Bertasius, and D. Crandall, ``Timerefine: Temporal grounding with time refining video llm,'' arXiv:2412.09601, 2024

  106. [114]

    H. Li, J. Chen, Z. Wei, S. Huang, T. Hui, J. Gao, X. Wei, and S. Liu, ``Llava-st: A multimodal large language model for fine-grained spatial-temporal understanding,'' arXiv:2501.08282, 2025

  107. [115]

    S. Chen, X. Lan, Y. Yuan, Z. Jie, and L. Ma, ``Timemarker: A versatile video-llm for long and short video understanding with superior temporal localization ability,'' arXiv:2411.18211, 2024

  108. [116]

    Q. Chen, S. Di, and W. Xie, ``Grounded multi-hop videoqa in long-form egocentric videos,'' in AAAI, 2025, pp. 2159--2167

  109. [117]

    M. Nie, D. Ding, C. Wang, Y. Guo, J. Han, H. Xu, and L. Zhang, ``Slowfocus: Enhancing fine-grained temporal understanding in video llm,'' in NeurIPS, 2024

  110. [118]

    Meinardus, A

    B. Meinardus, A. Batra, A. Rohrbach, and M. Rohrbach, ``The surprising effectiveness of multimodal large language models for video moment retrieval,'' arXiv:2406.18113, 2024

  111. [119]

    F. J. Fateh, U. Ahmed, H. Khan, M. Z. Zia, and Q.-H. Tran, ``Video llms for temporal reasoning in long videos,'' arXiv:2412.02930, 2024

  112. [120]

    Y. Wang, Y. Wang, P. Wu, J. Liang, D. Zhao, Y. Liu, and Z. Zheng, ``Efficient temporal extrapolation of multimodal large language models with temporal grounding bridge,'' in EMNLP, 2024, pp. 9972--9987

  113. [121]

    X. Li, B. Wang, G. Shi, C. Feng, and J. Teng, ``Mitigating the discrepancy between video and text temporal sequences: A time-perception enhanced video grounding method for llm,'' in COLING, 2025, pp. 9804--9813

  114. [122]

    Z. Yan, Z. Li, Y. He, C. Wang, K. Li, X. Li, X. Zeng, Z. Wang, Y. Wang, Y. Qiao et al., ``Task preference optimization: Improving multimodal large language models with vision task alignment,'' arXiv:2412.19326, 2024

  115. [123]

    J. Wang, Z. Liu, Y. Li, J. Ge, H. Xie, Y. Zhang et al., ``Spacevllm: Endowing multimodal large language model with spatio-temporal video grounding capability,'' arXiv:2503.13983, 2025

  116. [124]

    Z. Pang, M. Otani, and Y. Nakashima, ``Measure twice, cut once: Grasping video structures and event semantics with llms for video temporal localization,'' arXiv:2503.09027, 2025

  117. [125]

    Y. Wang, Z. Wang, B. Xu, Y. Du, K. Lin, Z. Xiao, Z. Yue, J. Ju, L. Zhang, D. Yang, X. Fang, Z. He, Z. Luo, W. Wang, J. Lin, J. Luan, and Q. Jin, ``Time-r1: Post-training large vision language model for temporal video grounding,'' arXiv:2503.13377, 2025

  118. [126]

    Zhao, G.-P

    H. Zhao, G.-P. Ji, R. Yan, H. Xiong, and Z. Li, ``Videoexpert: Augmented llm for temporal-sensitive video understanding,'' arXiv:2504.07519, 2025

  119. [127]

    X. Li, Z. Yan, D. Meng, L. Dong, X. Zeng, Y. He, Y. Wang, Y. Qiao, Y. Wang, and L. Wang, ``Videochat-r1: Enhancing spatio-temporal perception via reinforcement fine-tuning,'' arXiv:2504.06958, 2025

  120. [128]

    F. Luo, S. Lou, C. Chen, Z. Wang, C. Li, W. Shen, J. Guo, P. Li, M. Yan, J. Zhang et al., ``Museg: Reinforcing video temporal understanding via timestamp-aware multi-segment grounding,'' arXiv:2505.20715, 2025

  121. [129]

    Touvron, T

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi \`e re, N. Goyal, E. Hambro, F. Azhar et al., ``Llama: Open and efficient foundation language models,'' arXiv:2302.13971, 2023

  122. [130]

    Grauman, A

    K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu et al., ``Ego4d: Around the world in 3,000 hours of egocentric video,'' in CVPR, 2022, pp. 18\,995--19\,012

  123. [131]

    [Online]

    OpenAI, ``Gpt-4o,'' 2024. [Online]. Available: https://openai.com/index/hello-gpt-4o/

  124. [132]

    G. Team, P. Georgiev, V. I. Lei, R. Burnell, L. Bai, A. Gulati, G. Tanzer, D. Vincent, Z. Pan, S. Wang et al., ``Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context,'' arXiv:2403.05530, 2024

  125. [133]

    M. Maaz, H. Rasheed, S. Khan, and F. S. Khan, ``Video-chatgpt: Towards detailed video understanding via large vision and language models,'' in ACL, 2024

  126. [134]

    Reimers and I

    N. Reimers and I. Gurevych, ``Sentence-bert: Sentence embeddings using siamese bert-networks,'' in EMNLP, 2019, pp. 3982--3992

  127. [135]

    J. Chen, D. Zhu, X. Shen, X. Li, Z. Liu, P. Zhang, R. Krishnamoorthi, V. Chandra, Y. Xiong, and M. Elhoseiny, ``Minigpt-v2: large language model as a unified interface for vision-language multi-task learning,'' arXiv:2310.09478, 2023

  128. [136]

    A. Yang, B. Xiao, B. Wang, B. Zhang, C. Bian, C. Yin, C. Lv, D. Pan, D. Wang, D. Yan et al., ``Baichuan 2: Open large-scale language models,'' arXiv:2309.10305, 2023

  129. [137]

    K. C. Fraser and S. Kiritchenko, ``Examining gender and racial bias in large vision-language models using a novel dataset of parallel images,'' in EACL, 2024, pp. 690--713

  130. [138]

    Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu et al., ``Deepseekmath: Pushing the limits of mathematical reasoning in open language models,'' arXiv:2402.03300, 2024

  131. [139]

    Zellers, X

    R. Zellers, X. Lu, J. Hessel, Y. Yu, J. S. Park, J. Cao, A. Farhadi, and Y. Choi, ``Merlot: Multimodal neural script knowledge models,'' in NeurIPS, 2021, pp. 23\,634--23\,651

  132. [140]

    Y. Wang, Y. He, Y. Li, K. Li, J. Yu, X. Ma, X. Li, G. Chen, X. Chen, Y. Wang et al., ``Internvid: A large-scale video-text dataset for multimodal understanding and generation,'' arXiv:2307.06942, 2023

  133. [141]

    Zellers, J

    R. Zellers, J. Lu, X. Lu, Y. Yu, Y. Zhao, M. Salehi, A. Kusupati, J. Hessel, A. Farhadi, and Y. Choi, ``Merlot reserve: Neural script knowledge through vision and language and sound,'' in CVPR, 2022, pp. 16\,375--16\,387

  134. [142]

    Y. Zhai, S. Tong, X. Li, M. Cai, Q. Qu, Y. J. Lee, and Y. Ma, ``Investigating the catastrophic forgetting in multimodal large language models,'' arXiv:2309.10313, 2023

  135. [143]

    S. Chen, W. Jiang, W. Liu, and Y.-G. Jiang, ``Learning modality interaction for temporal sentence localization and event captioning in videos,'' in ECCV, 2020, pp. 333--351

  136. [144]

    Wang, Z.-J

    H. Wang, Z.-J. Zha, L. Li, D. Liu, and J. Luo, ``Structured multi-level interaction network for video moment localization via language query,'' in CVPR, 2021, pp. 7026--7035

  137. [145]

    J. Wang, L. Ma, and W. Jiang, ``Temporally grounding language queries in videos by contextual boundary-aware prediction,'' in AAAI, 2020, pp. 12\,168--12\,175

  138. [146]

    Chen, Y.-H

    Y.-W. Chen, Y.-H. Tsai, and M.-H. Yang, ``End-to-end multi-modal video temporal grounding,'' in NeurIPS, 2021, pp. 28\,442--28\,453

  139. [147]

    S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, and F. Wei, ``Beats: Audio pre-training with acoustic tokenizers,'' arXiv:2212.09058, 2022

  140. [148]

    Devlin, M.-W

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, ``Bert: Pre-training of deep bidirectional transformers for language understanding,'' in NAACL, 2019, pp. 4171--4186

  141. [149]

    J. Gao, C. Sun, Z. Yang, and R. Nevatia, ``Tall: Temporal activity localization via language query,'' in ICCV, 2017, pp. 5267--5275

  142. [150]

    L. Zhou, C. Xu, and J. Corso, ``Towards automatic learning of procedures from web instructional videos,'' in AAAI, 2018, pp. 7590--7598

  143. [151]

    G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta, ``Hollywood in homes: Crowdsourcing data collection for activity understanding,'' in ECCV, 2016, pp. 510--526

  144. [152]

    Caba Heilbron, V

    F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles, ``Activitynet: A large-scale video benchmark for human activity understanding,'' in CVPR, 2015, pp. 961--970

  145. [153]

    J. Xiao, X. Shang, A. Yao, and T.-S. Chua, ``Next-qa: Next phase of question-answering to explaining temporal actions,'' in CVPR, 2021, pp. 9777--9786

  146. [154]

    R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell, ``Natural language object retrieval,'' in CVPR, 2016, pp. 4555--4564

  147. [155]

    Fujita, T

    S. Fujita, T. Hirao, H. Kamigaito, M. Okumura, and M. Nagata, ``Soda: Story oriented dense video captioning evaluation framework,'' in ECCV, 2020, pp. 517--531

  148. [156]

    Banerjee and A

    S. Banerjee and A. Lavie, ``Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,'' in ACL Workshop, 2005, pp. 65--72

  149. [157]

    Vedantam, C

    R. Vedantam, C. Lawrence Zitnick, and D. Parikh, ``Cider: Consensus-based image description evaluation,'' in CVPR, 2015, pp. 4566--4575

  150. [158]

    W. Liu, T. Mei, Y. Zhang, C. Che, and J. Luo, ``Multi-task deep visual-semantic embedding for video thumbnail selection,'' in CVPR, 2015, pp. 3707--3715

  151. [159]

    Zhong, W

    Y. Zhong, W. Ji, J. Xiao, Y. Li, W. Deng, and T.-S. Chua, ``Video question answering: Datasets, algorithms and challenges,'' in EMNLP, 2022, pp. 6439--6455

  152. [160]

    X. Wang, Q. Si, J. Wu, S. Zhu, L. Cao, and L. Nie, ``Retake: Reducing temporal and knowledge redundancy for long video understanding,'' arXiv:2412.20504, 2024

  153. [161]

    Y. Li, C. Wang, and J. Jia, ``Llama-vid: An image is worth 2 tokens in large language models,'' in ECCV, 2024, pp. 323--340

  154. [162]

    Viertola, V

    I. Viertola, V. Iashin, and E. Rahtu, ``Temporally aligned audio for video with autoregression,'' in ICASSP, 2025, pp. 1--5

  155. [163]

    Yariv, I

    G. Yariv, I. Gat, S. Benaim, L. Wolf, I. Schwartz, and Y. Adi, ``Diverse and aligned audio-to-video generation via text-to-video model adaptation,'' in AAAI, 2024, pp. 6639--6647

  156. [164]

    Radford, J

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, ``Robust speech recognition via large-scale weak supervision,'' in ICML, 2023, pp. 28\,492--28\,518

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.