Pith. sign in

REVIEW 15 cited by

Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08085 v2 pith:AF4Z2JBR submitted 2024-06-12 cs.CV

classification cs.CV
keywords videoexistingunderstandingmodelsofflineonlinestreamsflash-vstream
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Benefiting from the advancements in large language models and cross-modal alignment, existing multi-modal video understanding methods have achieved prominent performance in offline scenario. However, online video streams, as one of the most common media forms in the real world, have seldom received attention. Compared to offline videos, the 'dynamic' nature of online video streams poses challenges for the direct application of existing models and introduces new problems, such as the storage of extremely long-term information, interaction between continuous visual content and 'asynchronous' user questions. Therefore, in this paper we present Flash-VStream, a video-language model that simulates the memory mechanism of human. Our model is able to process extremely long video streams in real-time and respond to user queries simultaneously. Compared to existing models, Flash-VStream achieves significant reductions in inference latency and VRAM consumption, which is intimately related to performing understanding of online streaming video. In addition, given that existing video understanding benchmarks predominantly concentrate on offline scenario, we propose VStream-QA, a novel question answering benchmark specifically designed for online video streaming understanding. Comparisons with popular existing methods on the proposed benchmark demonstrate the superiority of our method for such challenging setting. To verify the generalizability of our approach, we further evaluate it on existing video understanding benchmarks and achieves state-of-the-art performance in offline scenarios as well. All code, models, and datasets are available at the https://invinciblewyq.github.io/vstream-page/

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding

    cs.CV 2025-05 conditional novelty 7.0 of 10

    ScaleLong embeds four timescale question types into the same long videos, and evaluation of 23 MLLMs reveals a U-shaped accuracy curve across timescales.

  2. Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos

    cs.CV 2026-07 conditional novelty 6.5 of 10

    EgoMemo uses multi-scale temporal summaries, a knowledge graph, and visual archives to decide whether and when to intervene proactively on continuous egocentric video, setting baselines on the new EgoServe benchmark o...

  3. Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory

    cs.CV 2026-07 conditional novelty 6.5 of 10

    Dual global+latent states with hierarchical episodic merging enable reflexive, low-latency long-video agents that beat iterative reasoning baselines on accuracy and efficiency.

  4. GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience

    cs.CV 2026-08 conditional novelty 6.0 of 10

    GROVE shows that a streaming video memory stratified into four temporal scales, each with its own retrieval skill, improves both question answering and proactive assistance across five benchmarks.

  5. ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Training-free latent-object memory anchors let frozen Video-LLMs retain object histories under a tight token budget and improve streaming and long-video QA.

  6. ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding

    cs.AI 2026-07 conditional novelty 6.0 of 10

    ViSAGE builds entity-centered, self-correcting memories for long-form video understanding and reports state-of-the-art accuracy on M3-Bench and Video-MME-long.

  7. VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An open 4B video MLLM with inflated-3D ViT tokenization and adaptive streaming perception outperforms comparable open models on general, long-video, and streaming benchmarks while using fewer visual tokens.

  8. Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A hierarchical event memory with segment-tree proposals and a future prediction branch achieves state-of-the-art online video temporal grounding on TACoS, ActivityNet Captions, and MAD.

  9. Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    OpenBench, a new benchmark with categories semantically far from the COCO training space, shows that fine-tuning CLIP hurts open-vocabulary segmentation, and the proposed OVSNet method achieves state-of-the-art on bot...

  10. MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AVHaystacks is a new 3100-question benchmark for audio-visual QA across 500 videos, and the MAGNET multi-agent pipeline beats current baselines on it.

  11. Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Codec-guided sparse patch selection plus a lightweight speak/silent gate yields a 4B streaming VLM that is competitive on static tasks, stronger on video/spatial benchmarks, and much cheaper at inference.

  12. Vision-Language Memory for Spatial Reasoning

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A video-based vision-language model with 3D-aligned visual features and bounded dual memory achieves state-of-the-art scores on four spatial reasoning benchmarks.

  13. UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning

    cs.CV 2025-05 conditional novelty 5.0 of 10

    UniVG-R1 uses CoT supervised fine-tuning plus GRPO with difficulty-aware reweighting to make Qwen2-VL substantially better at multi-image, reasoning-based visual grounding.

  14. An Updated SynthPop Model for Microlensing Simulations I: Model Description & Evaluation

    astro-ph.GA 2026-03 unverdicted novelty 4.0 of 10

    An updated SynthPop model matches most bulge stellar and kinematic data but overpredicts optical microlensing event rates by about 20 percent near the galactic plane.

  15. Online Long-term Point Tracking in the Foundation Model Era

    cs.CV 2025-07 conditional novelty 2.0 of 10

    A frame-by-frame point tracker with spatial and context memory reaches accuracy comparable to offline trackers on seven video benchmarks, making online long-term point tracking feasible.

Pith tools