REVIEW 12 cited by
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This work proposes TimeChat, a time-sensitive multimodal large language model specifically designed for long video understanding. Our model incorporates two key architectural contributions: (1) a timestamp-aware frame encoder that binds visual content with the timestamp of each frame, and (2) a sliding video Q-Former that produces a video token sequence of varying lengths to accommodate videos of various durations. Additionally, we construct an instruction-tuning dataset, encompassing 6 tasks and a total of 125K instances, to further enhance TimeChat's instruction-following performance. Experiment results across various video understanding tasks, such as dense captioning, temporal grounding, and highlight detection, demonstrate TimeChat's strong zero-shot temporal localization and reasoning capabilities. For example, it achieves +9.2 F1 score and +2.8 CIDEr on YouCook2, +5.8 HIT@1 on QVHighlights, and +27.5 R@1 (IoU=0.5) on Charades-STA, compared to state-of-the-art video large language models, holding the potential to serve as a versatile video assistant for long-form video comprehension tasks and satisfy realistic user requirements.
Forward citations
Cited by 12 Pith papers
-
AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning
A video reasoning model learns per question whether to reason aloud or answer directly, improving accuracy by about 3 points over the best adaptive baseline while using about 23% fewer output tokens.
-
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
TimeLens2 shows that a compact video MLLM can localize multiple evidence intervals in long videos by training on verified interval labels and a Wasserstein-based time-distance reward.
-
TimePLE: Rethinking Temporal Representation for Video Temporal Grounding
TimePLE predicts a whole video interval as a joint distribution over a position-duration square, rather than predicting start and end separately, and reports higher mIoU across four VTG benchmarks.
-
Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing
Structured intermediate probing that separates modality-specific from modality-general signals lets privileged training modalities improve single-modality MLLM inference by large margins over naive multimodal training.
-
EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection
Separating temporal grounding from answer reasoning, plus low-confidence full-video re-reading, modestly improves long-video QA on Qwen3-VL across five benchmarks.
-
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences
Rule-reward training on controllable cross-video differences (Grounding + MCQ) improves Video MLLM local spatiotemporal evidence localization and transfers to general video QA benchmarks.
-
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
A training-free token compression method using semantic connected components in space and time keeps video understanding accuracy high even when retaining only 5-10% of visual tokens.
-
EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs
EASG-Bench is a 1,807-question egocentric video QA benchmark built from action scene graphs, and current video-LLMs score far below language-only models on temporal ordering questions.
-
RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
RICO refines image captions by reconstructing them into images with a text-to-image model and asking GPT-4o to fix discrepancies against the original, iteratively, with a DPO-distilled fast variant.
-
MTPChat: A Multimodal Time-Aware Persona Dataset for Conversational Agents
MTPChat adds explicit date stamps and synthetic earlier responses to multimodal persona dialogues, defines two temporal retrieval tasks, and reports modest gains from a gated fusion module.
-
MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding
A multimodal retrieval pipeline that projects video and audio into text and claims near-optimal context selection, with reported gains of up to 22.6% on Video-MME that rest on circular theory and unreleased data.
-
DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025
DIVE, an iterative question-decomposition system with intent estimation and object-centric video summarization, achieves 81.44% on CVRR-ES.
Discussion (0). Continue with ORCID to comment.