Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:37:08.992465Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 9 inbound Pith citation observations for arXiv:2506.23825.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:37:08.992465Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:03:35.208400Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z
85 of 85 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6887147d-f18e-4ce0-be0b-649e9ee97769 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f951860e-258b-40dd-9b3e-695ac42e79ff · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Self-calibrated clip for training-free open-vocabulary segmentation
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8c54936-2933-4df9-a915-0306862a540a · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Memory consolidation enables long-context video understanding
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1144657-4b06-4101-b70e-f10325e7f806 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Language models are few-shot learners
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed047457-67d2-493f-b1bd-4a5a5b54c0da · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams A Memory-Network Based Solution for Multivariate Time-Series Forecasting
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation caa6d757-081e-424c-9949-076661df0e62 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Distributed deep learning model for intelligent video surveillance systems with edge computing
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aafc111a-b9fa-4326-966c-e495ab223121 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Videollm-online: Online video large language model for streaming video
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3ad585c-95d6-4593-ae46-fc2760f3b7ac · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Sharegpt4video: Improving video understanding and generation with better captions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f9f6b6c-3478-4537-9fe5-05fe9a3a0293 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a8c5c172-dc6c-4e12-afaf-1b283327193c · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fcbabcd-d5a3-40ab-b2f5-a01ba03c21a5 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Instructblip: Towards general-purpose vision-language models with instruction tuning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 522ef2db-b67c-4aa7-924f-78c744c54477 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Flashattention-2: Faster attention with better paral- lelism and work partitioning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f8ab067d-b2d0-4b68-ace1-74339480aca4 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams An image is worth 16x16 words: Transformers for image recognition at scale
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f202760-4744-4c6a-aa31-8f2171aa0c05 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53b89593-713c-4fb7-b545-9c19be088905 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams A density-based algorithm for discovering clusters in large spatial databases with noise
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f0d1da68-38ef-429f-8509-457ba7bebf19 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 726d63d5-f4ff-4c22-bf44-8fd0f11dc009 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Temporal sentence grounding in streaming videos
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 49749f40-590f-4bee-9800-da523fe23e36 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Mist: Multi-modal iterative spatial- temporal transformer for long-form video question answering
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f38e2cb4-da5b-433f-92b1-b18c187df245 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Clip- adapter: Better vision-language models with feature adapters
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a7e01be7-e4e9-49f0-822a-ac95500b5f06 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Frameexit: Conditional early exiting for efficient video recognition
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9b2576bf-59fc-4109-8f36-de56f08fe307 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Dynamic neural networks: A survey
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 908cd346-dabb-4b80-93c1-c82eb7e70c04 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams A twofold siamese network for real-time object tracking
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e0f959ed-ecc4-46cc-b449-3d955ef8b133 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LoRA: Low-rank adaptation of large language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c2a9f090-f48f-4e98-b09c-0c5cd7f447a8 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Movienet: A holistic dataset for movie understanding
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 461b260e-ed3d-44a6-9ecc-b9003b49afea · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Chat-univi: Unified visual representation em- powers large language models with image and video under- standing
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 20d5ba49-83be-4a5b-9471-84ffaab9a37b · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Seed-bench: Benchmarking multimodal large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2bb99f58-390d-4cf3-a8b6-fd826d376847 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LLaVA-OneVision: Easy Visual Task Transfer
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c647090a-7381-4c06-bed0-c3a1c5d8f60c · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 391d44d2-f142-4fd4-9d13-ea37546e6ca8 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d3073307-a94a-44a6-b442-c23c61166907 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 82c186d8-45f2-428d-b58a-35d1f51e6d15 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Llama-vid: An image is worth 2 tokens in large language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 17555231-f16f-4331-b585-71124f0408f2 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Visual instruction tuning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1d17eb93-f65c-41fa-9216-1e97d9c997c0 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Improved baselines with visual instruction tuning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0f4f6773-ee34-41bd-bd60-e24713ee2a90 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe3eef6e-9ba6-4747-9420-16d2f1d9c9b1 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Learning quality-aware dynamic mem- ory for video object segmentation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 181b243a-8ba6-427b-a710-d3b95a117f09 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Universal segmentation at arbitrary granularity with language instruction
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9a7864cc-bbaa-41c0-bd7b-dca40e139fa4 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc6f9ea1-d404-4abd-b1e3-0c16b116cd10 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Soc: Semantic- assisted object cluster for referring video object segmentation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8f68580d-5ba6-40b9-a3d8-7caff45da478 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Multi-task deep learning for real-time 3d human pose estimation and action recognition
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 22555c9a-deb7-4a70-8ce1-2f4a2a885b9d · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation be611fb7-0e21-471b-8211-369fb2e96bde · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Video-chatgpt: Towards detailed video under- standing via large vision and language models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 41fb913b-96b3-4fec-9d33-77a80d89d812 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Some methods for classification and anal- ysis of multivariate observations
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f33954e-d803-4a93-a719-e53fc98c3706 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ba0da123-c123-4aa5-b457-02ca53d8da66 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Deepres: A deep learning-based video summarization strategy for resource- constrained industrial surveillance scenarios
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c70825d6-5f2e-4b0d-948d-316be74c920a · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Training language models to follow instructions with human feedback
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a26de3d5-160b-42bd-b8b9-9b830f97177c · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Streaming long video understanding with large language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4b17bef0-607e-4494-8825-a72a1012c341 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f2a3b187-e579-4d79-abc2-17de6c4e4e55 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Robovqa: Multimodal long-horizon reasoning for robotics
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 964afe38-b9f3-4395-b522-4bdd7d6198ae · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Moviechat: From dense token to sparse memory for long video understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 53fc9ae4-2eb8-4020-acdd-603aa9cd8429 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Roformer: Enhanced transformer with rotary position embedding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2563ac0-4c65-4138-859c-d37055ffaebf · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Tracking as online decision-making: Learning a policy from streaming videos with reinforcement learning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cfcecbf7-a784-46a9-b8e0-7978fc80d6fd · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Dynamic memory based attention network for sequential recommendation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6125cca0-ee24-438a-a6f3-d2b3b76defe7 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b37c7254-c536-42ff-82bb-7afa328b9445 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Kimi-VL Technical Report
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb49f620-fa9a-4472-9dfe-870c0cb1ac04 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LLaMA: Open and Efficient Foundation Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87ecba3e-14eb-4e4e-a8ff-8721d5970736 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a350224f-da1f-4b50-90cf-84aaac305b3e · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77574e63-3d4e-4b9b-adf3-d410f4e6881a · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LVBench: An Extreme Long Video Understanding Benchmark
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 242fa570-3d84-43c4-8f27-c01f37a67a4e · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f9bbaf5-76ae-4e83-8855-cd7253172285 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Adaptive focus for efficient video recognition
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f9a7a55a-0646-40da-a775-58ee96e5faa1 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Adafocus v2: End-to-end training of spatial dynamic net- works for video recognition
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 81c0ee5e-a38e-419d-a8b0-f15e86e45ecf · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Adafocusv3: On unified spatial-temporal dynamic video recognition
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bbdda88f-2070-44df-9029-f53700619c6f · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Hierarchical Memory for Long Video QA
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c31022a-e811-44a1-adfe-b975d579c6d7 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Ponder & Press: Advancing Visual GUI Agent towards General Computer Control
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 983783b2-0c3d-4687-88f8-c2bab012ce25 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Uni-adafocus: Spatial-temporal dynamic computation for video recognition
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 05e41abc-eb7f-4d65-90f7-facf52cbcefb · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Iterprime: Zero-shot referring image segmentation with iterative grad-cam refinement and primary word emphasis
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 108b516f-0500-4305-a3ed-f0802abfa1fa · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Sam2-love: Segment anything model 2 in language- aided audio-visual scenes
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9f00587a-b7bb-4c24-812d-f4a16f01545a · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Towards real-time multi-object tracking
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 66990da4-9ee6-4cdd-894d-9844412dd7d1 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Videollm-mod: Efficient video- language streaming with mixture-of-depths vision compu- tation
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f474a2e7-700c-40c7-beab-a09b604054f4 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Next-qa: Next phase of question-answering to explaining temporal actions
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 93c29c6b-51ce-4a07-9a67-99f57cf17501 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eb08b61-673f-44b1-8016-d638a7fe5d3d · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Fine-grained video captioning via graph-based multi- granularity interaction learning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b050451d-f53c-491f-950b-74967d66ffce · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Qwen2 Technical Report
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbe6b9ba-319a-4d92-9783-f9ee5860840b · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Language-aware vision transformer for referring segmentation
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1cf98fa4-906c-44af-af7c-e0594050c865 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Atp-llava: Adaptive token pruning for large vision language models
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d93939c1-ba4a-49c8-bdbf-69ed035ef2a7 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams V oco-llama: Towards vision compression with large language models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f40af02a-f349-43ba-b593-e05875441c11 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Self-chained image-language model for video localization and question answering
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5708aa3f-d747-4b72-9ae2-314e41a9b0a1 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3b095665-9420-4c13-8203-70c5b0f8f2b4 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Real-time action recognition with enhanced motion vector cnns
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fafdc095-d006-47ba-8474-889537b0e4ff · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Long Context Transfer from Language to Vision
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 113072f7-b149-48b4-98a8-8e9f4e1254e5 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52e61739-f8b5-43b7-a14c-62e6067bddfc · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams MLVU: Benchmarking Multi-task Long Video Understanding
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c8153d6-b46f-4322-b8f9-89626a8f23b0 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Streaming dense video captioning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation def39505-06cc-4ccf-baed-4dff0a300a59 · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams InstaRevive: One-Step Image Enhancement via Dynamic Score Matching
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e10025c3-8e4e-4ce7-8f23-a47b69e63cce · outbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Unresolved cited work
Reference 1000
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5d958aba-4e18-443b-bce7-01c2b8959940 · inbound
Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 011ffc0d-af10-4df5-b428-909d0d7ea02e · inbound
Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f3dff6f-4043-41f6-b592-79e133eaebbc · inbound
Mosaic: Cross-Modal Clustering for Efficient Video Understanding Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ce12786c-d594-42aa-8810-0d56971e0a6c · inbound
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 78041cb7-9628-4162-a2e3-c47795275266 · inbound
OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 44ceb307-4b36-4750-bc72-23d3b53ade81 · inbound
Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cf4b5ff5-efc3-4be8-997b-d36a18201cae · inbound
Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 54e3e141-b2f2-4945-be55-c0e5ffa8023f · inbound
MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 061c197b-c067-4d9a-bdf5-0e5971b2fd3a · inbound
Harnessing Streaming Video in the Wild Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.