Pith. sign in

Paper Citation Record · LEDGER

Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 50 inbound Pith citation observations for arXiv:2406.08085.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.08085 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 50 of 50 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:56.680449Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:00:08.182505Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 16ce91fa-49dc-4fb9-85aa-a081ec43c810 · inbound

LongVILA: Scaling Long-Context Visual Language Models for Long Videos cites this paper.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.463933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:2f307adaa0424216ed0ac4dea9267c4b0733dda4409b518cc9e56a6af00fde96

Observation fed152d1-acea-4a91-8536-678bbe328e50 · inbound

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance cites this paper.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:33:15.672665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:b58047323ee9191581487e5761d1f558265e0f7f664032981dda1fe875570a7f

Observation 80098db2-8922-4721-9090-bfaa8ada79d8 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 168

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.236233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:c8c1a6444cac0f0bc48ac8a4605dc87e4f92326fba7383c8b3ccb490fe6bc892

Observation 739c46fe-1d13-4e13-b3db-19f83f83e070 · inbound

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning cites this paper.

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:56.680449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:56.680449Z digest=sha256:98730d1f9973c74497b9610d33c308bdf9c0d4097b2e4fcca29094e354f45eae

Observation 9c3d523b-1f4c-4892-95fa-71444da65956 · inbound

LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval cites this paper.

LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:31:40.813023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T14:26:59.015559Z digest=sha256:08eabf6053a138e02933638ba7c70657c470775f830c95a52668cb64d769f5f2

Observation f614c2f9-c1ee-4bd9-92c9-241b1d2d38bb · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:55:40.390683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:51247d023a8e1427f4c7cc256c51f2013e6d0bd2b313df027b6d3ad06260c357

Observation c6df2d43-09d9-4618-97e4-0c342718411e · inbound

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding cites this paper.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:46.869024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:46.869024Z digest=sha256:5368588f035974b1f6e9d8c41ff4878b1231a67846b4f5cd6c17ebe9b0218648

Observation 5c7ac51c-e03f-4c20-be75-737f1f5e4b4e · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.419480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.419480Z digest=sha256:79041f5308a410042f3b47e839fc36993ba9ed2688f4b89f2c155cd66e08542c

Observation fe2be44c-1e16-44a6-ad2d-f8e543fb0667 · inbound

Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation cites this paper.

Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T23:50:23.560907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:50:23.560907Z digest=sha256:990990e2a0bb18ca069e75b1eccb92e192812f37ac4299ac31e16f7ceb5c1bc1

Observation 8e73ad35-a823-4d03-a8a7-bb0f21cdc78c · inbound

Online Long-term Point Tracking in the Foundation Model Era cites this paper.

Online Long-term Point Tracking in the Foundation Model Era Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:06:18.952812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:06:18.952812Z digest=sha256:3d1a304f336d170b1eed9993f45635dbbd405f40a030d6822a58a010cc9b1d1b

Observation 2ba91c3e-102d-4de4-a5c1-c99f9b4b88de · inbound

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding cites this paper.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.777349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.777349Z digest=sha256:92f7ac4785ce0e4b468bd8112333115d4d90970300cdcdec04846d04c477674e

Observation 946eac70-1f6c-4433-aad0-c82da8c1daea · inbound

Vision-Language Memory for Spatial Reasoning cites this paper.

Vision-Language Memory for Spatial Reasoning Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:36.429995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:36.429995Z digest=sha256:087a38869bdb284e5896ac1da2f6a112e65205c15442b0bb12ec2742cee6a1a9

Observation 14008280-962a-4104-a8cc-86e1bb31ca92 · inbound

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance? cites this paper.

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance? Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:39:06.167340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T05:36:09.208754Z digest=sha256:bc8adb7f63c16b7ffb3107a8a5e092a13bb0eefee893fb897293174ab858a225

Observation 394e8922-e627-4b86-b997-907e05b8b432 · inbound

Streaming Video Instruction Tuning cites this paper.

Streaming Video Instruction Tuning Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.792961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T19:44:11.032898Z digest=sha256:ac0c3884b0d28c4ec3e7c3e50a890dde2b682a76d6940cacf70f641b374a1766

Observation 22a96c46-ffa6-459c-a1f9-b95b727ad2cb · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:53.864829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:3ab68f307d66dd67cd10e855607c527ef989b71e9d9468f6ab3290a19f6dde8d

Observation 0bf7a826-66ac-4eb0-a31b-6e0dc62ccb89 · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.505455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:410bc0b9da2117702be158392c041a5eb29c31582a2ee786a4a4aea94a2b585a

Observation 5bc8b24f-ac61-4f4b-a48e-e571ca4389d6 · inbound

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents cites this paper.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:10:13.107496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:1597f054418197b6f2bd5052d24e058d9c78ddba834911243df5379db661e28b

Observation 104221c3-5d21-4f4e-a341-9336cef722da · inbound

An Updated SynthPop Model for Microlensing Simulations I: Model Description & Evaluation cites this paper.

An Updated SynthPop Model for Microlensing Simulations I: Model Description & Evaluation Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T22:25:30.596294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:25:30.596294Z digest=sha256:51974fb0d53367a42c65d21d9b13afdb01151fb4245bdbc31dcf8754eafd273d

Observation df9ac8cf-62b0-49a1-b05c-90746584bb94 · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.448097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:623aef466ad6242d80adae9177689d1678b39e0c57e1628748d8ca79fcf57acd

Observation 2583ddc4-2dc7-4abc-af36-05ebbeaae521 · inbound

VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models cites this paper.

VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:25:58.031174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:39:30.731688Z digest=sha256:0c951c0116ccfdb1d4d03bb420ab66ba3c4ee9b1c2aa18065ff707a1121f5938

Observation 8cb884b5-c0e4-4c76-b400-7aa8cb670cf2 · inbound

StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding cites this paper.

StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.436845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:19:41.543165Z digest=sha256:c67a41021e248c33c33040b02e960e9481dbb9d20df961bc7eea74d24de8211e

Observation 266777e1-7e61-40fb-aea8-674cb00bbeca · inbound

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning cites this paper.

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.788325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T06:51:52.861981Z digest=sha256:b40d81c0259dfd100d9569e2dd0fd38739f7b04e4ab772e834cd8cfa8cb2f08c

Observation 13efc00b-3054-4d1d-8fd7-92eba1529fe7 · inbound

Don't Pause! Every prediction matters in a streaming video cites this paper.

Don't Pause! Every prediction matters in a streaming video Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:18.019791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T04:32:01.379605Z digest=sha256:f028f61ed8e067b548a1af7b5dcfd138ef2740497cb179a9149d8893d71d57b6

Observation 09c9c0bb-1fff-461f-aa8e-83141dc06e75 · inbound

Decouple and Cache: KV Cache Construction for Streaming Video Understanding cites this paper.

Decouple and Cache: KV Cache Construction for Streaming Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:31:03.562532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T14:47:54.917408Z digest=sha256:43b48d69d1e811ea66930600f042a50a8cdab3d606a7718c7b00b7af6cd2530d

Observation 5cdcc589-c2b1-4b8e-8faa-0b46fd2d4312 · inbound

From Priors to Perception: Grounding Video-LLMs in Physical Reality cites this paper.

From Priors to Perception: Grounding Video-LLMs in Physical Reality Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:08.334591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T17:41:23.233366Z digest=sha256:31fc310c6e12878333300bc63f1b0f77e59469f5621c56ff2a63997bb6206f71

Observation 212f5093-1618-4c58-a2ad-0068af94c72a · inbound

LATERN: Test-Time Context-Aware Explainable Video Anomaly Detection cites this paper.

LATERN: Test-Time Context-Aware Explainable Video Anomaly Detection Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T14:25:47.016973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T21:19:12.655706Z digest=sha256:3c2afbbb8563bc1b83590691562e8a585d5f3443bcfd864bcea7b9712a5f3217

Observation eb91be19-68ab-41e9-ba79-cff4ed9f0561 · inbound

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning cites this paper.

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:24.798248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T15:12:00.408851Z digest=sha256:db43c2b792e51bfa8587c1645f5a46fedcca63c6f2f945c8b9a6ae5a1d7dec3e

Observation 0c263db3-57e5-4683-92be-a303128f8bbc · inbound

OProver: A Unified Framework for Agentic Formal Theorem Proving cites this paper.

OProver: A Unified Framework for Agentic Formal Theorem Proving Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:48:23.506925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T14:43:46.517807Z digest=sha256:3ce27e6b38a41931704f7f721d5b0ee8d81c1bceca773308eda4d12de5059054

Observation 96f99809-c799-4cf6-8451-346d8e4bcfc1 · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:38:19.169529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T13:36:44.071188Z digest=sha256:6895cae7e8a77bc87733e72957cbe23d5e847428503071b5da2a4a62741ddd6c

Observation aaf4586a-8910-4338-80b9-98681305850c · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:19:20.326643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-04T01:11:42.073993Z digest=sha256:ed356371510490d6c05b0896b8acbb5e2356fc80f7de4ac16f45d6a7737cee4f

Observation 6743be0a-b7b7-40fd-9707-80343302ea06 · inbound

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams cites this paper.

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:51.028063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T18:24:57.881644Z digest=sha256:1901917c5e784cff8218c87cdf19407fa70c1236c67166e28e9163e64a574b1c

Observation 2b18fe52-a846-4236-b189-fc792d0a4fb2 · inbound

Linear Scaling Video VLMs for Long Video Understanding cites this paper.

Linear Scaling Video VLMs for Long Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:02:46.258016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T23:00:11.246232Z digest=sha256:1fa11f48663fbf5d8f56bdf5e24fda6b38007057d3a2e487adfbbfa911f1f681

Observation 7356d735-315d-4c30-af27-76028d7db25c · inbound

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation cites this paper.

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:27.065722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T10:54:02.188634Z digest=sha256:e6e8a5eedfd54704392f66000d447e559583e4b183fc37429de884b9b02b20d6

Observation cb654ab0-e384-4fd2-9d22-52721b12820f · inbound

OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs cites this paper.

OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:27.100649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T11:02:07.122615Z digest=sha256:da81ac9f3e19dc8ec070237189f5d7b98f35a40248413d62947291267e52af4b

Observation f17c9624-dcc6-4b32-9028-2268977d50fc · inbound

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding cites this paper.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.826114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:cf6be776a9e42980ce972489dc34799eaafaf06f8b89cfeb8636b42506bc638b

Observation ace06fdb-019e-44c5-aecf-8cefd999604a · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.497535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:c31494d21f8a1cf69a203606d91b8f9e8dba16bdeb018484d89829ba26c9534e

Observation 8b545684-c27d-47da-8667-b0cb0d11a8eb · inbound

Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur? cites this paper.

Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur? Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:30.647830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T16:52:22.811857Z digest=sha256:1adeb18ac1bc10abaabe48864b3a18d25c8f3e5c3bfbcca20e28bdf16032ea70

Observation 8b069b59-ddbf-4ffd-8919-b3520863a203 · inbound

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams cites this paper.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.143340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:7381b344e40c6ded1ff499ce7526dda1751464fc9a4458177f48b2c18c76b83d

Observation 587d4ed6-2de8-4cf1-aff5-54b31c752bf1 · inbound

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference cites this paper.

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:29.905981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T18:17:53.013043Z digest=sha256:39b7f1c3aff8429d4968a14c6e8a8ebc81075ce9dab505e7de14d6ef769c23ea

Observation 25d70545-f7d4-4401-ad83-bd7f31670d73 · inbound

How Well Can Your Video Model Remember? Measuring Memory-Budget Trade-offs in Long Video Understanding cites this paper.

How Well Can Your Video Model Remember? Measuring Memory-Budget Trade-offs in Long Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:39:04.710461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T21:51:04.050833Z digest=sha256:751506c18854534b1ec4a074882e22dee72515d215569e185c7523582f130c33

Observation 16e9037b-f4fe-44d4-8fa8-a2fee29d1c6c · inbound

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding cites this paper.

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:39:58.343073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T00:19:26.153682Z digest=sha256:b8367e4d65ff03468f40931f6a0f79436cc9e5706ce0dab2f3137703570ec69e

Observation bf1fbb9e-037d-40dd-9ba6-a6f5309f6bd5 · inbound

Towards a Dynamic and Fixed-budget Memory Bank for Efficient Streaming Video Understanding cites this paper.

Towards a Dynamic and Fixed-budget Memory Bank for Efficient Streaming Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:00:08.184260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T20:54:49.319252Z digest=sha256:77d5e9d25e6e4b406dd0ab6987ba144018a5636356bc73f2a125bb3bb5338fe4

Observation 97a1906b-585e-4bfb-9a68-31c654f7fd22 · inbound

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory cites this paper.

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-11T06:35:35.951554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T06:35:35.951554Z digest=sha256:39dcc7620d2ca5dabf8e9066565dcda14c74d8d195383bd711e04a1a143bb3a6

Observation 8226f807-a8ea-4980-a197-0d07ec3918a5 · inbound

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos cites this paper.

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-14T05:01:06.200663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:01:06.200663Z digest=sha256:1075049bc20bf96bb4027f28386c824ba35e26bf6ed39bd24b405eb25fe1461c

Observation 7e5440df-9915-40ca-8717-78f8973438bf · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.589936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.589936Z digest=sha256:6af7c9857416d6103896939be4da1d541a32fa01c9a789c39545411e73d35740

Observation d760ee98-5e90-420d-9bca-e53aecc7cd85 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.884368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.884368Z digest=sha256:3082bd718b23657196bd323ad7da62ff467ef2e27eeae1b61685d24ca281b556

Observation 1b043da0-8aaa-49d4-99f7-8bb3ff0b33e3 · inbound

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding cites this paper.

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T03:21:48.375332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:21:48.375332Z digest=sha256:4ee6cb03b7bb3eed9f60b3d77991f06fe30925b061c5a65e00d8dde54701389e

Observation 9cea1237-dc36-4c46-8a46-6741e1a482ff · inbound

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding cites this paper.

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T00:45:21.931265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:45:21.931265Z digest=sha256:a8e5797155458344cdb136679da4a3a185c90f84520d7af9af6c62d1ea786309

Observation 259dbeea-6a1e-41a5-89ad-a27f7e1e4aaf · inbound

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience cites this paper.

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T08:12:47.458517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:12:47.458517Z digest=sha256:ffb8d537a00a84183fee7857c025b59b0946808d036f8dd499fbd764b0fbc9ce

Observation 499c13e1-ce69-46d3-89ec-54ff6b1c3754 · inbound

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience cites this paper.

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:15.929069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:15.929069Z digest=sha256:6b3c28e8f3b3bedff63223522f275152604c66c9eaacccae059ffdab3b47701b