Pith. sign in

REVIEW 5 major objections 6 minor 3 cited by

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?

T0 review · 5 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read GLIMPSE is a new benchmark built to test whether vision-language models can reason across an entire video rather than answer from a few frames, and it reports that the best model tested, GPT-o3, scores 66.43% while humans score 94.82%.

desk verdict Useful new benchmark, but the paper's central claim that all questions require full-video temporal reasoning is not yet proven, and Section 4.4 partly contradicts it. read the letter →

arxiv 2507.09491 v1 pith:F452E74K submitted 2025-07-13 cs.CV

classification cs.CV
keywords videounderstandingbenchmarklargevision-languagemodelstemporalreasoningvisual-centricquestionansweringQAevaluationdynamicsLVLM
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

GLIMPSE is a benchmark built to test whether large vision-language models can actually reason across an entire video, rather than answer from a handful of frames. It contains 3,269 videos and 4,342 multiple-choice questions, written by human annotators, that the authors argue cannot be solved by glancing at still images or by relying on language priors alone. On a 550-question human subset, volunteers scored 94.82%, while the strongest model tested, GPT-o3, reached 66.43%, and most open-source video models landed between roughly 40% and 52%. The benchmark's central claim is that this gap reflects missing true video reasoning in current models, not just frame-level perception.

What carries the argument

The load-bearing mechanism is the benchmark construction procedure: human annotators watch each full video and write questions that require cross-frame analysis; a manual quality review removes questions answerable from a single frame; GPT-4o converts open-ended annotations into multiple-choice form; and yes/no temporal questions are paired with reversed-order versions so a model is credited only if it answers both directions correctly. This paired-question design, together with the 11 visual-centric categories (trajectory analysis, temporal reasoning, quantitative estimation, event recognition, reverse event inference, scene context awareness, velocity estimation, cinematic dynamics, forensic authenticity analysis, robotics evaluation, and multi-object interaction), is what lets GLIMPSE separate genuine video reasoning from frame scanning.

What would settle it

Give a model or a human only one randomly selected frame from each GLIMPSE video and measure accuracy on the full question set. If a single-frame baseline answers a substantial share of questions above chance, especially in trajectory, event, or temporal categories, the claim that GLIMPSE questions require full-video reasoning would be contradicted.

Watch

Extended reading notes

Core claim

The central discovery is a measured performance gap that the benchmark is designed to make visible: when questions require tracking trajectories, comparing event order, estimating counts from motion, judging camera motion, or spotting generated video anomalies, state-of-the-art vision-language models perform far below human level. The paper's evidence is that the strongest evaluated model, GPT-o3, reaches 66.43% average accuracy, Gemini 1.5 Pro reaches 56.98%, GPT-4o reaches 53.80%, and the best open-source video model, Qwen2-VL, reaches 52.47%, while human volunteers reached 94.82%. The authors take this as evidence that current models have not yet achieved the full-video reasoning the benchmark targets, and that image-only models scoring 37.48% or lower help confirm the questions are not answerable from one static view.

Load-bearing premise

The load-bearing assumption is that the annotators' manual review reliably removed every question that could be answered from one frame or a few frames, and the paper offers no objective check of that on the full benchmark.

Editorial extensions

If this is right

  • The strongest tested closed-source model still leaves roughly one in three GLIMPSE questions wrong, so claims that current vision-language models understand video should be treated as partial.
  • Image-only models score near or below chance on temporal reasoning, so GLIMPSE provides evidence that its tasks genuinely distinguish video understanding from still-image understanding.
  • Video-specialized models beat image models on temporal categories, but their overall accuracy remains below 53%, showing that video training helps but does not close the human-model gap.
  • Temporal reasoning accuracy declines for videos longer than 30 seconds, indicating that long-video temporal understanding is a distinct bottleneck.
  • Requiring correct answers on reversed yes/no pairs substantially lowers reported accuracy, showing that simple frame scanning or answer-order bias is not enough to solve the benchmark.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One extension the paper leaves implicit is that GLIMPSE could serve as an audit tool for frame sampling: a model's score depends heavily on how many frames it is given, as the paper's own frame-count experiment on quantitative estimation shows.
  • A testable next step is a text-only control that presents questions and answer options without any video, which would quantify how much language priors alone contribute to correct answers.
  • If the measured gap is real, training objectives that reward temporal coherence across frames or architectures that preserve long-range motion features should move GLIMPSE scores, making the benchmark a possible training target rather than only an evaluation set.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. The paper introduces GLIMPSE, a video question-answering benchmark of 3,269 videos and 4,342 human-written multiple-choice questions across 11 visual-centric categories (trajectory analysis, temporal reasoning, quantitative estimation, forensics, robotics, multi-object interaction, and others). The central claim is that, unlike prior video benchmarks, GLIMPSE questions cannot be answered by scanning a few frames or relying on text alone; they require reasoning over the full video. Evaluations of commercial and open-source LVLMs show a large human-model gap: human volunteers reach 94.82% accuracy on a 550-question subset, while the best model, GPT-o3, reaches only 66.43%, and most open-source video models score between 39% and 53%. The paper concludes that current LVLMs lack genuine full-video temporal reasoning and proposes GLIMPSE as a diagnostic benchmark for such reasoning.

Significance. If the central claim is verified, GLIMPSE is a valuable and much-needed benchmark: it targets temporal and dynamic reasoning skills under-represented in existing single-frame-answerable video benchmarks, uses manually curated videos across diverse domains, introduces bidirectional yes/no pairs to mitigate position bias, and evaluates a broad range of models under a consistent protocol. The benchmark and code are publicly released, and the reported numbers are internally consistent. The main caveat is that the premise that all questions require full-video reasoning rests entirely on a subjective manual review; if this premise is weakened, the headline human-model gap could be inflated by construction artifacts. The paper would be substantially strengthened by a frame-count ablation across all categories and a human single-frame baseline to objectively substantiate the full-video requirement.

major comments (5)
  1. [Abstract and §3.2] The claim that GLIMPSE questions 'cannot be answered by scanning selected frames or relying on text alone' is not objectively verified. The quality review in §3.2 is the only support, but the paper reports no removal statistics, no inter-annotator agreement, and no explicit test of single-frame answerability. The single-frame image-LVLM results in Table 1 use one randomly sampled frame per video, which can miss the informative key frame, so they do not establish that a question cannot be solved from a few well-chosen frames. Please add a frame-count ablation across all question categories (e.g., 1, 2, 4, 8, 16 frames) and/or a human single-frame baseline; without this, the 94.82% versus 66.43% gap cannot be attributed to full-video reasoning.
  2. [§4.4] Section 4.4 states that Quantitative Estimation 'requires detailed snapshots of individual frames, as these provide data points for counting or assessing quantities.' This explicitly identifies a snapshot-answerable category, which contradicts the abstract's claim that the benchmark's questions cannot be answered by scanning selected frames. The authors should either restrict the claim to the subset of categories proven to require full-video reasoning or add a per-category analysis of frame-answerability and adjust the narrative accordingly.
  3. [§4.1 Human Evaluation] The human baseline is computed from a 550-question subset (50 per category) answered by five student volunteers, with no reported variance, no inter-annotator agreement, and no justification for calling the volunteers 'experts.' Models are evaluated on all 4,342 questions, so the comparison of 94.82% (human subset) to 66.43% (full benchmark) is only valid if the subset is representative and the human measurement is stable. Please report per-category question counts, confidence intervals for both human and model accuracies, and a human agreement statistic (e.g., Fleiss' kappa).
  4. [§3.2 and Appendix A] The GPT-4o-based reformulation of open-ended annotations into multiple-choice questions, and the generation of bidirectional yes/no pairs, may introduce systematic bias. The prompt in Appendix A instructs the model to make questions 'significantly harder and more nuanced' and to add 'realistic but irrelevant details,' which could produce ambiguous or artificially difficult items that are not well-calibrated to the video content. The paper does not report any validation that the reformatted questions preserve the original intent or that the reversed pairs are logical inverses. A human agreement study on the reformatted items, and a comparison between original and reformatted question performance, would address this risk.
  5. [§4.1 Experiment Setup] Closed-source models are evaluated with fixed-interval sampling capped at 16 frames, while open-source video models use their default frame counts. Since the benchmark claims to assess full-video reasoning, the 16-frame cap may differentially disadvantage commercial models. The frame-count analysis in §4.4 covers only Quantitative Estimation; please report sensitivity results for closed-source models with different caps (e.g., 8, 16, 32 frames) across all categories, or justify the cap as sufficient for the benchmark's question types.
minor comments (6)
  1. [§3.2] The text states that video length was controlled to be 'between 20 seconds and 2 minutes,' but Figure 4 and Table 2 include 0-10s and 10-20s bins. Please reconcile this inconsistency.
  2. [§2 and §3.2] There are repeated typos in model names and text: 'LLA V A' should be 'LLaVA', 'GA V A' should be 'GLaVA' (or intended model name), 'open sourece' appears in the Introduction, and 'V elocity' appears in §3.2.
  3. [§4.2 and §4.3] References to 'Table 4' should be to 'Table 1' (e.g., 'Based on the results in Table 4' and 'The test settings were the same as in Table 4' in §4.3).
  4. [§4.3] The caption of Table 2 says 'across 12 categories,' but GLIMPSE has 11 categories; please correct.
  5. [Abstract and §3.2] The abstract says 'over 4,342' questions while §3.2 says 'a total of 4342'; please use a consistent number.
  6. [§4.1] The Random baseline computation for bidirectional temporal-reasoning pairs is not described; please clarify how pairs are treated in the random baseline.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: GLIMPSE model accuracies are external measurements, and the full-video requirement is enforced during curation rather than imported as a self-cited theorem or fitted parameter.

full rationale

GLIMPSE is a benchmark paper; its central empirical claims are the measured accuracies in Table 1 and the human evaluation in Section 4.1. These numbers are not derived from the benchmark's own definitions or from any parameter fitted to the data, so the headline gap (94.82% human vs. 66.43% GPT-o3) is an independent measurement. The property that questions 'cannot be answered by scanning selected frames' is introduced as a curation criterion (Section 3.2: 'We also ensured that each question required understanding of the entire video rather than a single frame. For example, questions like "What is the weather in the video?" could be answered with just one frame, so similar questions were filtered out during the review process.') rather than obtained as an output of the evaluation; it is a definitional design goal, not a fitted prediction. The paper's self-citations (e.g., Ye et al. 2023, Wang et al. 2024b, Zhou et al. 2024) appear in related-work or baseline contexts and are not load-bearing for the benchmark's conclusions. No uniqueness theorem or prior result by the authors is invoked to force the benchmark's choices. One caveat worth recording is not circularity but internal tension: Section 4.4 states that Quantitative Estimation 'requires detailed snapshots of individual frames, as these provide data points for counting or assessing quantities,' and the quality-review statistics (how many questions were removed, inter-annotator agreement) are not reported; this bears on whether every question is genuinely full-video, but it does not make the measured LVLM-human gap a consequence of the paper's own definitions. Likewise, using a randomly sampled single frame for image LVLMs does not prove that no single key frame suffices, but that is a benchmarking-design strength question, not a circularity in the derivation chain. Overall, the paper's results are self-contained external measurements, so the circularity score is 0.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the quality of the human annotation and filtering process rather than on fitted parameters or invented entities. The only hand-chosen evaluation parameter is the frame sampling for commercial models. The main axioms are domain assumptions about the reliability of GPT-4o-based conversion, the small human evaluation, and the representativeness of video sources.

free parameters (1)
  • Closed-source model frame sampling = sample_frequency=50, max_frame_num=16
    Chosen by hand for all commercial models; limits the amount of video a model sees and could affect performance on long videos.
assumptions (4)
  • domain assumption GPT-4o API conversion preserves the correct answer and introduces no systematic bias.
    Section 3.2 and Appendix A use a prompt that deliberately makes questions harder and embeds the correct answer as a subtle hint; errors here would corrupt the benchmark ground truth.
  • domain assumption Human volunteer answers are an accurate ground truth and upper bound.
    Section 4.1: five students answered 550 questions after watching once, with no inter-annotator agreement reported; a small sample may not represent true human performance.
  • ad hoc to paper The manual quality review successfully excludes all single-frame-answerable questions.
    Quality Review (Section 3.2) states such questions are filtered out, but no independent verification (e.g., single-frame human baseline) is provided.
  • domain assumption Videos from selected public sources are representative of dynamic video understanding.
    Video Collection (Section 3.2) draws from Pexels, Pixabay, Ego4D, Panda-70m, ShareGPTVideo, PushT, and Cable Routing; coverage of real-world temporal scenarios is assumed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?." pith.science (2026). https://pith.science/paper/F452E74K

@misc{pith2026250709491,
  author       = {Pith},
  title        = {Pith review of: GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/F452E74K}},
  note         = {Machine review of arXiv:2507.09491}
}
read the original abstract

Existing video benchmarks often resemble image-based benchmarks, with question types like "What actions does the person perform throughout the video?" or "What color is the woman's dress in the video?" For these, models can often answer by scanning just a few key frames, without deep temporal reasoning. This limits our ability to assess whether large vision-language models (LVLMs) can truly think with videos rather than perform superficial frame-level analysis. To address this, we introduce GLIMPSE, a benchmark specifically designed to evaluate whether LVLMs can genuinely think with videos. Unlike prior benchmarks, GLIMPSE emphasizes comprehensive video understanding beyond static image cues. It consists of 3,269 videos and over 4,342 highly visual-centric questions across 11 categories, including Trajectory Analysis, Temporal Reasoning, and Forensics Detection. All questions are carefully crafted by human annotators and require watching the entire video and reasoning over full video context-this is what we mean by thinking with video. These questions cannot be answered by scanning selected frames or relying on text alone. In human evaluations, GLIMPSE achieves 94.82% accuracy, but current LVLMs face significant challenges. Even the best-performing model, GPT-o3, reaches only 66.43%, highlighting that LVLMs still struggle to move beyond surface-level reasoning to truly think with videos.

Figures

Figures reproduced from arXiv: 2507.09491 by the authors.

Figure 1
Figure 1. LVLMs without reasoning ability struggle [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. GLIMPSE categorizes visual-centric questions in video data into 11 distinct types, with representative examples from each category illustrated in the figure. Human-annotated questions and answers are reformatted into multiple-choice format for LVLMs. To reduce bias, yes/no questions are presented bidirectionally, requiring correct answers in both directions to be considered accurate. person moves closer to or farthe… view at source ↗
Figure 3
Figure 3. Distribution of categories in the GLIMPSE benchmark. Our benchmark covers 4 key domains and 11 detailed visual-centric question types. laying the foundation for video-based multimodal dialogue. Subsequent research introduced improve￾ments (Zhang et al., 2023; Jin et al., 2023; Ren et al., 2024; Song et al., 2024; Liu et al., 2023a), including the addition of audio modalities, joint training on images and videos, and… view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Video length distribution, with the horizontal [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: Model performance comparison in Scene Con [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Model performance on quantitative estimation [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 8
Figure 8. Figure 8: The failure cases of GPT-4o in GLIMPSE. It can be observed that GPT-4o has difficulty accurately understanding quantity-related questions in the video. Additionally, when answering temporal reasoning questions, it makes mistakes when the question is reversed. ing state…
Figure 9
Figure 9. Figure 9: Dataset cases 1. 14 [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: Dataset cases 2. 15 [PITH_FULL_IMAGE:figures/full_fig_p015_10.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Low-Cost Test-Time Adaptation for Robust Video Editing

    cs.CV 2025-07 reject novelty 5.0 of 10

    Vid-TTA proposes to adapt video editing UNets per test video via motion-aware masked autoencoding and prompt perturbation, with claimed but unquantified improvements.

  2. A Multimodal Deep Learning Framework for Early Diagnosis of Liver Cancer via Optimized BiLSTM-AM-VMD Architecture

    cs.LG 2025-09 reject novelty 3.0 of 10

    The paper claims a BiLSTM-AM-VMD model achieves AUC 0.963 for early HCC diagnosis, but the evidence is undermined by contradictory dataset descriptions and missing artifacts.

  3. Multi-Modal Machine Learning Framework for Predicting Early Recurrence of Brain Tumors Using MRI and Clinical Biomarkers

    cs.LG 2025-09 reject novelty 3.0 of 10

    XGBoost combining MRI radiomics and clinical biomarkers reportedly reaches C-index 0.782 for early brain tumor recurrence, but the paper's methods describe a liver-cancer cohort and no evaluation of its claimed tempor...

Reference graph

Works this paper leans on

21 extracted references · 1 canonical work pages · cited by 3 Pith papers

  1. [3]

    arXiv preprint arXiv:2311.14906

    Autoeval-video: An automatic bench- mark for assessing large vision language models in open-ended video question answering. arXiv preprint arXiv:2311.14906. Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, and 1 others

  2. [4]

    arXiv preprint arXiv:2406.07476

    Videol- lama 2: Advancing spatial-temporal modeling and audio understanding in video-llms. arXiv preprint arXiv:2406.07476. Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song

  3. [5]

    arXiv preprint arXiv:2405.21075

    Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. arXiv preprint arXiv:2405.21075. Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, and 1 others

  4. [6]

    arXiv preprint arXiv:2311.08046

    Chat-univi: Unified vi- sual representation empowers large language models with image and video understanding. arXiv preprint arXiv:2311.08046. Peng Jin, Ryuichi Takanobu, Wancai Zhang, Xiaochun Cao, and Li Yuan

  5. [7]

    ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models

    Vilma: A zero-shot bench- mark for linguistic and temporal grounding in video- language models. arXiv preprint arXiv:2311.07022. Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yix- iao Ge, and Ying Shan. 2023a. Seed-bench: Bench- marking multimodal llms with generative compre- hension. arXiv preprint arXiv:2307.16125. KunChang Li, Yinan He, Yi Wang, Yizhuo...

  6. [8]

    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22195–22206

    Mvbench: A com- prehensive multi-modal video understanding bench- mark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22195–22206. Shuailin Li, Yuang Zhang, Yucheng Zhao, Qiuyue Wang, Fan Jia, Yingfei Liu, and Tiancai Wang. 2023c. Vlm-eval: A general evaluation on video large lan- guage models. arXiv preprint ...

  7. [9]

    arXiv preprint arXiv:2311.10122

    Video-llava: Learn- ing united visual representation by alignment before projection. arXiv preprint arXiv:2311.10122. Haogeng Liu, Qihang Fan, Tingkai Liu, Linjie Yang, Yunzhe Tao, Huaibo Huang, Ran He, and Hongxia Yang. 2023a. Video-teller: Enhancing cross-modal generation with fusion and decoupling. arXiv preprint arXiv:2310.04991. Haotian Liu, Chunyuan...

  8. [10]

    arXiv preprint arXiv:2306.05424

    Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424. Karttikeya Mangalam, Raiymbek Akshulakov, and Ji- tendra Malik

Show all 21 references
  1. [11]

    arXiv preprint arXiv:2311.16103

    Video- bench: A comprehensive benchmark and toolkit for evaluating video-based large language models. arXiv preprint arXiv:2311.16103. OpenAI

  2. [12]

    arXiv preprint arXiv:2307.01952

    Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952. Huaizhi Qu, Xinyu Zhao, Jie Peng, Kwonjoon Lee, Behzad Dariush, and Tianlong Chen. Uq-merge: Uncertainty guided multimodal large language model merging. Shuhuai Ren, L...

  3. [13]

    arXiv preprint arXiv:2312.11805

    Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805. Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, and 1 others

  4. [14]

    arXiv preprint arXiv:2403.05530

    Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530. Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar,...

  5. [15]

    arXiv preprint arXiv:2302.13971

    Llama: Open and effi- cient foundation language models. arXiv preprint arXiv:2302.13971. 11 Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhi- hao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, and 1 others. 2024a. Qwen2- vl: Enhancing vision-language model’s...

  6. [16]

    arXiv preprint arXiv:2404.16994

    Pllava: Parameter-free llava extension from images to videos for video dense captioning. arXiv preprint arXiv:2404.16994. Jihan Yang, Shusheng Yang, Anjali W Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie

  7. [17]

    arXiv preprint arXiv:2304.14178

    mplug-owl: Modularization empowers large lan- guage models with multimodality. arXiv preprint arXiv:2304.14178. Qinghao Ye, Haiyang Xu, Jiabo Ye, Ming Yan, An- wen Hu, Haowei Liu, Qi Qian, Ji Zhang, and Fei Huang

  8. [19]

    arXiv preprint arXiv:2306.02858

    Video- llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858. Yuanhan Zhang, Bo Li, haotian Liu, Yong jae Lee, Liangke Gui, Di Fu, Jiashi Feng, Ziwei Liu, and Chunyuan Li

  9. [20]

    arXiv preprint arXiv:2405.14622

    Cali- brated self-rewarding vision language models. arXiv preprint arXiv:2405.14622. Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny

  10. [21]

    {original_question}

    Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592. 12 A Prompt for Data Format Conversion In this section, we list the prompts used during the data conversion process, where the entire data format transforma...

  11. [2023]

    arXiv preprint arXiv:2308.12966

    Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966. Mu Cai, Reuben Tan, Jianrui Zhang, Bocheng Zou, Kai Zhang, Feng Yao, Fangrui Zhu, Jing Gu, Yiwu Zhong, Yuzhang Shang, and 1 others

  12. [2024]

    arXiv preprint arXiv:2410.10818

    Tempo- ralbench: Benchmarking fine-grained temporal un- derstanding for multimodal video models. arXiv preprint arXiv:2410.10818. Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Bin Lin, Zhenyu Tang, Li Yuan, Yu Qiao, Dahua Lin, ...

  13. [2025]

    arXiv preprint arXiv:2502.12081

    Unhack- able temporal rewarding for scalable video mllms. arXiv preprint arXiv:2502.12081. Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yuet- ing Zhuang, and Dacheng Tao

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.