Pith. sign in

REVIEW 42 cited by

StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.03628 v1 pith:DIQG5CLE submitted 2024-11-06 cs.CV cs.AI

StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

classification cs.CV cs.AI
keywords videomllmsunderstandingstreamingstreamingbenchcapabilitiescomprehensionbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The rapid development of Multimodal Large Language Models (MLLMs) has expanded their capabilities from image comprehension to video understanding. However, most of these MLLMs focus primarily on offline video comprehension, necessitating extensive processing of all video frames before any queries can be made. This presents a significant gap compared to the human ability to watch, listen, think, and respond to streaming inputs in real time, highlighting the limitations of current MLLMs. In this paper, we introduce StreamingBench, the first comprehensive benchmark designed to evaluate the streaming video understanding capabilities of MLLMs. StreamingBench assesses three core aspects of streaming video understanding: (1) real-time visual understanding, (2) omni-source understanding, and (3) contextual understanding. The benchmark consists of 18 tasks, featuring 900 videos and 4,500 human-curated QA pairs. Each video features five questions presented at different time points to simulate a continuous streaming scenario. We conduct experiments on StreamingBench with 13 open-source and proprietary MLLMs and find that even the most advanced proprietary MLLMs like Gemini 1.5 Pro and GPT-4o perform significantly below human-level streaming video understanding capabilities. We hope our work can facilitate further advancements for MLLMs, empowering them to approach human-level video comprehension and interaction in more realistic scenarios.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 42 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning

    cs.CV 2026-08 conditional novelty 7.0

    A video reasoning model learns per question whether to reason aloud or answer directly, improving accuracy by about 3 points over the best adaptive baseline while using about 23% fewer output tokens.

  2. VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance

    cs.CV 2026-07 conditional novelty 7.0

    VIABench provides 761 long-form egocentric videos from blind individuals with 14,526 annotations across three assistance tasks, and shows current multimodal LLMs achieve best overall scores below 30.

  3. EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding

    cs.CV 2026-06 unverdicted novelty 7.0

    EgoSAT is the first benchmark unifying retrospective, online, and prospective reasoning tasks in egocentric streaming video to evaluate VLMs, revealing struggles with temporal modeling and mis-calibration.

  4. X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding

    cs.CV 2026-06 unverdicted novelty 7.0

    X-Stream benchmark shows SOTA MLLMs score ~50% on concurrent multi-stream tasks and lack proactive ability, using a dual-verification pipeline to avoid single-stream bias.

  5. X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding

    cs.CV 2026-06 unverdicted novelty 7.0

    X-Stream benchmark shows state-of-the-art MLLMs achieve only about 50% on multi-stream video tasks and exhibit poor proactive ability.

  6. EGOSTREAM: A Diagnostic Benchmark for Streaming Episodic Memory in Egocentric Vision

    cs.CV 2026-05 unverdicted novelty 7.0

    Egostream introduces a diagnostic benchmark that expands 2,250 questions into 8,528 recall-conditioned evaluations to measure streaming episodic memory performance across detail, spatial, temporal, event, social, caus...

  7. An Efficient Streaming Video Understanding Framework with Agentic Control

    cs.CV 2026-05 unverdicted novelty 7.0

    R3-Streaming uses cascaded control with age-aware memory forgetting and TB-GRPO reinforcement learning to reach SOTA scores of 57.92 on OVO-Bench and 76.36 on StreamingBench with 95-96% fewer visual tokens.

  8. Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction

    cs.CV 2026-05 conditional novelty 7.0

    Omni-DuplexEval creates a new benchmark and LLM-as-a-Judge framework for real-time duplex omni-modal interaction, revealing that current models score below 40% overall and struggle especially with proactive responses.

  9. Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding

    cs.CV 2026-05 unverdicted novelty 7.0

    SAVEMem improves streaming video understanding scores by adding semantic awareness to memory compression and query-adaptive retrieval without any model training.

  10. Don't Pause! Every prediction matters in a streaming video

    cs.CV 2026-04 unverdicted novelty 7.0

    SPOT-Bench tests real-time streaming video perception with timeliness metrics, exposing limitations in current models and introducing AsynKV as an improved baseline.

  11. OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning

    cs.CV 2026-04 unverdicted novelty 7.0

    OASIS organizes streaming video into hierarchical events and retrieves memory on-demand via intent-driven refinement to improve long-horizon accuracy and compositional reasoning with bounded token costs.

  12. Online Reasoning Video Object Segmentation

    cs.CV 2026-04 unverdicted novelty 7.0

    The work introduces the ORVOS task, the ORVOSB benchmark with causal annotations across 210 videos, and a baseline using updated prompts plus a temporal token reservoir.

  13. VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models

    cs.CV 2026-04 unverdicted novelty 7.0

    VSAS-Bench offers temporally dense annotations and synchronous/asynchronous protocols to evaluate streaming VLMs on timeliness, consistency, accuracy, and latency trade-offs, showing that adapted conventional VLMs can...

  14. StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos

    cs.CV 2025-12 unverdicted novelty 7.0

    StreamGaze is a new benchmark and QA generation pipeline that measures how well MLLMs leverage gaze trajectories for temporal reasoning and proactive intention prediction in streaming egocentric videos.

  15. WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

    cs.CV 2025-02 unverdicted novelty 7.0

    WorldSense provides the first benchmark requiring synergistic audio-video-text understanding on 1,662 real-world videos and 3,172 QA pairs, where the best current multimodal LLM reaches only 65.1% accuracy.

  16. Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos

    cs.CV 2026-07 conditional novelty 6.5

    EgoMemo uses multi-scale temporal summaries, a knowledge graph, and visual archives to decide whether and when to intervene proactively on continuous egocentric video, setting baselines on the new EgoServe benchmark o...

  17. ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

    cs.CV 2026-07 conditional novelty 6.0

    A training-free memory framework that anchors streaming video memory to latent objects discovered from frozen Video-LLM features, improving streaming QA accuracy while cutting memory and latency.

  18. ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

    cs.CV 2026-07 conditional novelty 6.0

    Training-free latent-object memory anchors let frozen Video-LLMs retain object histories under a tight token budget and improve streaming and long-video QA.

  19. QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding

    cs.CV 2026-07 accept novelty 6.0

    QSVideo reformulates questions into structured queries, ranks frames by object-action-location relevance plus diversity, and applies temporal strategies to boost VLM accuracy under tight frame budgets.

  20. GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video

    cs.CV 2026-07 conditional novelty 6.0

    Current MLLMs can deliver procedural instructions in streaming video but systematically fail at real-time error detection and corrective coaching on the new GuideMe benchmark.

  21. MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding

    cs.CV 2026-07 unverdicted novelty 6.0

    MedStreamBench integrates 22 medical datasets into 5,419 QA instances across retrospective, present, future, and proactive temporal settings to evaluate streaming and proactive medical video understanding.

  22. ProtoKV: Streaming Video Understanding under Delayed Query with Summary-State Memory

    cs.CV 2026-06 unverdicted novelty 6.0

    ProtoKV maintains a fixed-capacity summary state for far history in streaming video, improving accuracy by up to 12.5 points in long-delay query scenarios compared to token-retention methods.

  23. LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs

    cs.DC 2026-06 unverdicted novelty 6.0

    LiveServe exposes audio playback and barge-in signals to the scheduler and KV manager, lowering P90 audio TTFP by 1.55x on average and raising completed-request throughput by 1.15x on two Omni-LMs.

  24. Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur?

    cs.CV 2026-06 unverdicted novelty 6.0

    Introduces Ego-MC-Bench benchmark and Ego-CoMist synthetic dataset showing that fine-tuning video LLMs on proactive mistake corrections improves performance especially for smaller models.

  25. Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding

    cs.CV 2026-06 unverdicted novelty 6.0

    LyraV uses FDTC and SToP for per-frame incremental decoding to reach 98.29% video synchrony at 3.89 FPS while preserving general understanding.

  26. MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention

    cs.CV 2026-06 unverdicted novelty 6.0

    MOSS-Video-Preview introduces a cross-attention architecture and synthesized real-time QA data to enable continuous perception, answer revision, and faster inference in video-language models compared to decoder-only designs.

  27. StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering

    cs.CV 2026-05 unverdicted novelty 6.0

    StreamOV proposes evidence-guided long-short term memory and a hidden-state-driven trigger for efficient online audio-visual reasoning in streaming videos, along with the SOVBench benchmark for multi-turn evaluation.

  28. An Efficient Streaming Video Understanding Framework with Agentic Control

    cs.CV 2026-05 unverdicted novelty 6.0

    R3-Streaming uses cascaded control, age-aware memory forgetting, and TB-GRPO reinforcement learning to reach SOTA scores on streaming video benchmarks while cutting visual token usage by 95-96%.

  29. Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction

    cs.CV 2026-05 unverdicted novelty 6.0

    Omni-DuplexEval provides a new benchmark and automatic evaluation method for real-time duplex omni-modal interaction, showing state-of-the-art models reach only 39.6% overall and 20% on proactive reminders.

  30. Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding

    cs.CV 2026-05 unverdicted novelty 6.0

    Response-G1 uses query-guided scene graphs, memory retrieval, and augmented prompting to improve when Video-LLMs decide to respond during streaming videos.

  31. SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance

    cs.CV 2026-03 accept novelty 6.0

    Streaming multi-point counting on 406 videos with three trajectory metrics reveals large human-model gaps in spatial-temporal state maintenance, worst on periodic events.

  32. Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously

    cs.CV 2026-03 conditional novelty 6.0

    A 7B video model that generates intermediate text thoughts during playback, before the query arrives, improves streaming-video QA accuracy while keeping query-time latency near real-time.

  33. LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval

    cs.CV 2025-05 unverdicted novelty 6.0

    LiveVLM introduces VSB and PaR to compress and retrieve KV cache in streaming video LLMs, enabling LLaVA-OneVision to reach SOTA accuracy among training-free query-agnostic and training-based online models.

  34. Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

    cs.CV 2026-07 conditional novelty 5.0

    Codec-guided sparse patch selection plus a lightweight speak/silent gate yields a 4B streaming VLM that is competitive on static tasks, stronger on video/spatial benchmarks, and much cheaper at inference.

  35. MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering

    cs.CV 2026-05 conditional novelty 5.0

    MuKV adds multi-grained KV cache compression at patch-frame-segment levels plus semi-hierarchical retrieval to raise accuracy and cut memory in long video question-answering.

  36. Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding

    cs.CV 2026-05 unverdicted novelty 5.0

    Response-G1 uses query-guided scene graph generation, memory retrieval, and retrieval-augmented prompting to improve proactive response timing in streaming video understanding.

  37. Decouple and Cache: KV Cache Construction for Streaming Video Understanding

    cs.CV 2026-05 unverdicted novelty 5.0

    DSCache decouples cumulative past and instant KV caches with position-agnostic encoding to adapt offline VideoVLLMs to streaming video, delivering 2.5% average accuracy gains on QA benchmarks.

  38. LiveVLN: Breaking the Stop-and-Go Loop in Vision-Language Navigation

    cs.RO 2026-04 unverdicted novelty 5.0

    LiveVLN enables smoother vision-language navigation by overlapping action execution with ongoing observation processing, preserving benchmark scores while cutting real-world waiting time by up to 77.7 percent.

  39. Existence of small semi-vortex solutions for the cubic nonlinear Schr\"{o}dinger system with Rashba type Spin-Orbit coupling on $\mathbb{R}^2$

    math.AP 2026-04 unverdicted novelty 5.0

    Existence of small semi-vortex solutions for the Rashba SOC cubic NLS system on R^2 is proved via energy minimization under small mass constraint.

  40. Existence of small semi-vortex solutions for the cubic nonlinear Schr\"{o}dinger system with Rashba type Spin-Orbit coupling on $\mathbb{R}^2$

    math.AP 2026-04 unverdicted novelty 5.0

    Small semi-vortex and ground-state solutions of the cubic NLS system with Rashba SOC on R² exist as energy minimizers under small mass, via concentration-compactness.

  41. Seed1.8 Model Card: Towards Generalized Real-World Agency

    cs.AI 2026-03 unverdicted novelty 5.0

    Seed1.8 is a new foundation model that adds unified agentic capabilities for search, code execution, and GUI interaction to existing LLM and vision strengths.

  42. Seed1.5-VL Technical Report

    cs.CV 2025-05 unverdicted novelty 4.0

    Seed1.5-VL is a compact multimodal model that sets new records on dozens of vision-language benchmarks and outperforms prior systems on agent-style tasks.