Pith. sign in

Paper Citation Record · LEDGER

Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 50 inbound Pith citation observations for arXiv:2406.08085.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.08085 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 50 of 50 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:56.680449Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:00:08.182505Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 16ce91fa-49dc-4fb9-85aa-a081ec43c810 · inbound

LongVILA: Scaling Long-Context Visual Language Models for Long Videos cites this paper.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.463933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:b7131d93fd3291b5517cd5c91cc82e4f09840630d58cd16a39e3af19ae33e9f1

Observation fed152d1-acea-4a91-8536-678bbe328e50 · inbound

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance cites this paper.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:33:15.672665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:b271a3d78e1c5519f0edd2163147f61e03eb5e2151dc0bb63fbd5139bcd63cff

Observation 80098db2-8922-4721-9090-bfaa8ada79d8 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 168

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.236233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:39e0ff984b6823c94e93208be09f6ef6bb247e8003e9620140a48973769c9cdb

Observation 739c46fe-1d13-4e13-b3db-19f83f83e070 · inbound

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning cites this paper.

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:56.680449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:56.680449Z digest=sha256:98730d1f9973c74497b9610d33c308bdf9c0d4097b2e4fcca29094e354f45eae

Observation 9c3d523b-1f4c-4892-95fa-71444da65956 · inbound

LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval cites this paper.

LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:31:40.813023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T14:26:59.015559Z digest=sha256:fa5fbdc0ff3c421b1789b7cee538a396f02fa465c940484971616f567cbca536

Observation f614c2f9-c1ee-4bd9-92c9-241b1d2d38bb · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:55:40.390683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:66d585ed45c54a35bc5daac1d36773f8a0417811c8987a1f583f8be6ef78e00e

Observation c6df2d43-09d9-4618-97e4-0c342718411e · inbound

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding cites this paper.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:46.869024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:46.869024Z digest=sha256:5368588f035974b1f6e9d8c41ff4878b1231a67846b4f5cd6c17ebe9b0218648

Observation 5c7ac51c-e03f-4c20-be75-737f1f5e4b4e · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.419480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.419480Z digest=sha256:cd9250ada6b6c604c76d9d69d5242e44f05e6bb30696c950ec2bd72b0756d772

Observation fe2be44c-1e16-44a6-ad2d-f8e543fb0667 · inbound

Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation cites this paper.

Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T23:50:23.560907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:50:23.560907Z digest=sha256:990990e2a0bb18ca069e75b1eccb92e192812f37ac4299ac31e16f7ceb5c1bc1

Observation 8e73ad35-a823-4d03-a8a7-bb0f21cdc78c · inbound

Online Long-term Point Tracking in the Foundation Model Era cites this paper.

Online Long-term Point Tracking in the Foundation Model Era Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:06:18.952812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:06:18.952812Z digest=sha256:3d1a304f336d170b1eed9993f45635dbbd405f40a030d6822a58a010cc9b1d1b

Observation 2ba91c3e-102d-4de4-a5c1-c99f9b4b88de · inbound

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding cites this paper.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.777349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.777349Z digest=sha256:92f7ac4785ce0e4b468bd8112333115d4d90970300cdcdec04846d04c477674e

Observation 946eac70-1f6c-4433-aad0-c82da8c1daea · inbound

Vision-Language Memory for Spatial Reasoning cites this paper.

Vision-Language Memory for Spatial Reasoning Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:36.429995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:36.429995Z digest=sha256:087a38869bdb284e5896ac1da2f6a112e65205c15442b0bb12ec2742cee6a1a9

Observation 14008280-962a-4104-a8cc-86e1bb31ca92 · inbound

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance? cites this paper.

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance? Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:39:06.167340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T05:36:09.208754Z digest=sha256:5fe581a7d801e86446b2f3e2b7908cbb3b9849c6631632b6909d65b14976f5e5

Observation 394e8922-e627-4b86-b997-907e05b8b432 · inbound

Streaming Video Instruction Tuning cites this paper.

Streaming Video Instruction Tuning Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.792961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T19:44:11.032898Z digest=sha256:d384b01f66bafe08418c5574abdd55996670c8e87e0aff568c7a0d7dd5cbe315

Observation 22a96c46-ffa6-459c-a1f9-b95b727ad2cb · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:53.864829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:a16fa74cf62320f69fb2944e0dfc47e8d4f5dcba9d589c678d0d9a95d1892fdb

Observation 0bf7a826-66ac-4eb0-a31b-6e0dc62ccb89 · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.505455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:56508331e02c8d2b4093d4767a264f71f80111aff45e704eed3a118e78a82b3b

Observation 5bc8b24f-ac61-4f4b-a48e-e571ca4389d6 · inbound

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents cites this paper.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:10:13.107496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:f7fb744c9d7a4b7c214edf4cbc1e0ab7fa9ccd3bc97e11d08556f46e12cab46e

Observation 104221c3-5d21-4f4e-a341-9336cef722da · inbound

An Updated SynthPop Model for Microlensing Simulations I: Model Description & Evaluation cites this paper.

An Updated SynthPop Model for Microlensing Simulations I: Model Description & Evaluation Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T22:25:30.596294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:25:30.596294Z digest=sha256:51974fb0d53367a42c65d21d9b13afdb01151fb4245bdbc31dcf8754eafd273d

Observation df9ac8cf-62b0-49a1-b05c-90746584bb94 · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.448097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:ed2e40d6992fe798e71ead35c7b641703d13cdd76c04893d47dad1cdf68eb499

Observation 2583ddc4-2dc7-4abc-af36-05ebbeaae521 · inbound

VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models cites this paper.

VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:25:58.031174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:39:30.731688Z digest=sha256:dbf2936d99ded12b8698988799b719d0b0c941be97cca610609be978b50568fc

Observation 8cb884b5-c0e4-4c76-b400-7aa8cb670cf2 · inbound

StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding cites this paper.

StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.436845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:19:41.543165Z digest=sha256:e80502d0853f7c1b34aa59d6f8b2ce76bec8d5e7407d431db6db536292b7a8e5

Observation 266777e1-7e61-40fb-aea8-674cb00bbeca · inbound

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning cites this paper.

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.788325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T06:51:52.861981Z digest=sha256:2ced9cfdd53d9087f652c265db3f27db64ab150e7cdcccc840635ab909b525fa

Observation 13efc00b-3054-4d1d-8fd7-92eba1529fe7 · inbound

Don't Pause! Every prediction matters in a streaming video cites this paper.

Don't Pause! Every prediction matters in a streaming video Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:18.019791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T04:32:01.379605Z digest=sha256:3fab515443186d7d074506158e7376f3faf9feabdafcd3cd6d3b0840f6c202ff

Observation 09c9c0bb-1fff-461f-aa8e-83141dc06e75 · inbound

Decouple and Cache: KV Cache Construction for Streaming Video Understanding cites this paper.

Decouple and Cache: KV Cache Construction for Streaming Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:31:03.562532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T14:47:54.917408Z digest=sha256:ffb7a3b46c8f16884bb2e0fc6c51ce14a1a6d12b9d5c50176486df4e0c2e0914

Observation 5cdcc589-c2b1-4b8e-8faa-0b46fd2d4312 · inbound

From Priors to Perception: Grounding Video-LLMs in Physical Reality cites this paper.

From Priors to Perception: Grounding Video-LLMs in Physical Reality Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:08.334591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T17:41:23.233366Z digest=sha256:1cfadb137287186b1bfc8e67f8a71b60b7d25f923b8aca3eabfe7006b76311a0

Observation 212f5093-1618-4c58-a2ad-0068af94c72a · inbound

LATERN: Test-Time Context-Aware Explainable Video Anomaly Detection cites this paper.

LATERN: Test-Time Context-Aware Explainable Video Anomaly Detection Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T14:25:47.016973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:19:12.655706Z digest=sha256:9c73ac76d945f78df1b1940c718d60b7196b81d3a04640ece6dbb393d8d8083f

Observation eb91be19-68ab-41e9-ba79-cff4ed9f0561 · inbound

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning cites this paper.

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:24.798248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T15:12:00.408851Z digest=sha256:7354dfde5951f158447a964c177d895451085a8c7d0500bf2b9f14503bc542a2

Observation 0c263db3-57e5-4683-92be-a303128f8bbc · inbound

OProver: A Unified Framework for Agentic Formal Theorem Proving cites this paper.

OProver: A Unified Framework for Agentic Formal Theorem Proving Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:48:23.506925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T14:43:46.517807Z digest=sha256:68fe403957a4d41c659c9065a469fb49771c35fbef94451fffdca36ffe1dae32

Observation 96f99809-c799-4cf6-8451-346d8e4bcfc1 · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:38:19.169529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T13:36:44.071188Z digest=sha256:a33be0ca838560f0cb32a4a6bbb5fe385c5000160441e82c33fc8a135a04801a

Observation aaf4586a-8910-4338-80b9-98681305850c · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:19:20.326643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-04T01:11:42.073993Z digest=sha256:0b9aa52bdcf37de59f9a51d87b49acc9f836737002feb9134874498b64c61166

Observation 6743be0a-b7b7-40fd-9707-80343302ea06 · inbound

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams cites this paper.

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:51.028063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T18:24:57.881644Z digest=sha256:897a50e9aaa1d61ddbd64f2bae02b01a8f3b9e3542205a244e1217361633a85b

Observation 2b18fe52-a846-4236-b189-fc792d0a4fb2 · inbound

Linear Scaling Video VLMs for Long Video Understanding cites this paper.

Linear Scaling Video VLMs for Long Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:02:46.258016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T23:00:11.246232Z digest=sha256:2dee22ac482bec615c56c3b32a692c322635937cbb00f0f89371251d1bacb96c

Observation 7356d735-315d-4c30-af27-76028d7db25c · inbound

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation cites this paper.

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:27.065722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T10:54:02.188634Z digest=sha256:97955bc5583a9b27f04760547c9783cc08771fca37fa50a564c323d54405873a

Observation cb654ab0-e384-4fd2-9d22-52721b12820f · inbound

OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs cites this paper.

OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:27.100649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T11:02:07.122615Z digest=sha256:b1ed534215b1606607611169e58a7cd96b74b23af5370bbfc8e75b4cd0a2bf0c

Observation f17c9624-dcc6-4b32-9028-2268977d50fc · inbound

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding cites this paper.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.826114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:3f6739a3fc9696f745b67891f9a3050335f56c3e17d662bbeca1b9477c6bfee6

Observation ace06fdb-019e-44c5-aecf-8cefd999604a · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.497535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:241537ff9d4ea1c23cf8b834d8696702e9de1658df7f650d380c32ca878ca4cf

Observation 8b545684-c27d-47da-8667-b0cb0d11a8eb · inbound

Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur? cites this paper.

Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur? Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:30.647830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:52:22.811857Z digest=sha256:98fa7c27ba0b44010c0a93b2d2b9d9dabc58e792469ee40ac664edbddff2db0e

Observation 8b069b59-ddbf-4ffd-8919-b3520863a203 · inbound

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams cites this paper.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.143340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:512ba781eadc905c659d497f6f7ea53b5936d743127a6a91c3efa6ed28f823f9

Observation 587d4ed6-2de8-4cf1-aff5-54b31c752bf1 · inbound

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference cites this paper.

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:29.905981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T18:17:53.013043Z digest=sha256:938e4b7feffbb55bff81ca312afe396a24c32ee801a48381f32c9ae6b7741de6

Observation 25d70545-f7d4-4401-ad83-bd7f31670d73 · inbound

How Well Can Your Video Model Remember? Measuring Memory-Budget Trade-offs in Long Video Understanding cites this paper.

How Well Can Your Video Model Remember? Measuring Memory-Budget Trade-offs in Long Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:39:04.710461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T21:51:04.050833Z digest=sha256:72b13311be5d85ffd2544c525260fcd6eda2a5644f62624f5067538085c90e4d

Observation 16e9037b-f4fe-44d4-8fa8-a2fee29d1c6c · inbound

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding cites this paper.

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:39:58.343073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T00:19:26.153682Z digest=sha256:1d9f34e3bc574498beb148c8dad31137649f8f41db6326538e4adea60cf232ab

Observation bf1fbb9e-037d-40dd-9ba6-a6f5309f6bd5 · inbound

Towards a Dynamic and Fixed-budget Memory Bank for Efficient Streaming Video Understanding cites this paper.

Towards a Dynamic and Fixed-budget Memory Bank for Efficient Streaming Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:00:08.184260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T20:54:49.319252Z digest=sha256:6f5a3ea9bda889262453dcbccb32c645d926dea345732dd60a1abc246df6b3b3

Observation 97a1906b-585e-4bfb-9a68-31c654f7fd22 · inbound

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory cites this paper.

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-11T06:35:35.951554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T06:35:35.951554Z digest=sha256:39dcc7620d2ca5dabf8e9066565dcda14c74d8d195383bd711e04a1a143bb3a6

Observation 8226f807-a8ea-4980-a197-0d07ec3918a5 · inbound

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos cites this paper.

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-14T05:01:06.200663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:01:06.200663Z digest=sha256:1075049bc20bf96bb4027f28386c824ba35e26bf6ed39bd24b405eb25fe1461c

Observation 7e5440df-9915-40ca-8717-78f8973438bf · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.589936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.589936Z digest=sha256:6af7c9857416d6103896939be4da1d541a32fa01c9a789c39545411e73d35740

Observation d760ee98-5e90-420d-9bca-e53aecc7cd85 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.884368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.884368Z digest=sha256:3082bd718b23657196bd323ad7da62ff467ef2e27eeae1b61685d24ca281b556

Observation 1b043da0-8aaa-49d4-99f7-8bb3ff0b33e3 · inbound

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding cites this paper.

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T03:21:48.375332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:21:48.375332Z digest=sha256:4ee6cb03b7bb3eed9f60b3d77991f06fe30925b061c5a65e00d8dde54701389e

Observation 9cea1237-dc36-4c46-8a46-6741e1a482ff · inbound

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding cites this paper.

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T00:45:21.931265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:45:21.931265Z digest=sha256:a8e5797155458344cdb136679da4a3a185c90f84520d7af9af6c62d1ea786309

Observation 259dbeea-6a1e-41a5-89ad-a27f7e1e4aaf · inbound

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience cites this paper.

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T08:12:47.458517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:12:47.458517Z digest=sha256:13f2c31cca7467f3295a09f15368d6e7366ab20f0e457e34618f136679841108

Observation 499c13e1-ce69-46d3-89ec-54ff6b1c3754 · inbound

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience cites this paper.

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:15.929069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:15.929069Z digest=sha256:e0358eaa6da0b65f6fcdff65e20f4b0ca2124227a6378587376a9da2ea991ebe