Pith. sign in

Paper Citation Record · LEDGER

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

As of 20 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 11 inbound Pith citation observations for arXiv:2506.23825.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23825 v2

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:37:08.992465Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:14:01.995575Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact2
  • verified fuzzy54
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6887147d-f18e-4ce0-be0b-649e9ee97769 · outbound

This paper cites GPT-4 Technical Report.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:01.869344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:01.869344Z digest=sha256:d48a5177596c0318c0170089100e32b334bbb012a39567685747dea47c951ee4

Observation f951860e-258b-40dd-9b3e-695ac42e79ff · outbound

This paper cites Self-calibrated clip for training-free open-vocabulary segmentation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Self-calibrated clip for training-free open-vocabulary segmentation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:01.948138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:01.948138Z digest=sha256:0f2ffd9250b1be65a6ced8123c09e636ac8abf5090d5c317829aeb48b14339fc

Observation e8c54936-2933-4df9-a915-0306862a540a · outbound

This paper cites Memory consolidation enables long-context video understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Memory consolidation enables long-context video understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:02.070349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:02.070349Z digest=sha256:6e3627e86a0194a6a51e3da2a84433aebbbe0808b99ffd0dce6fd36e6002b7d0

Observation f1144657-4b06-4101-b70e-f10325e7f806 · outbound

This paper cites Language models are few-shot learners.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Language models are few-shot learners

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:02.199304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:02.199304Z digest=sha256:2d010ff4d4b63925f5deb72f431cac8e87990173da460aa9724460f63b691a6a

Observation ed047457-67d2-493f-b1bd-4a5a5b54c0da · outbound

This paper cites A Memory-Network Based Solution for Multivariate Time-Series Forecasting.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams A Memory-Network Based Solution for Multivariate Time-Series Forecasting

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:37:09.602810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:02.378276Z digest=sha256:ca854bcc2521870c37582763305218d2c5e484879cb0b0fa384c1295567eabef

Observation caa6d757-081e-424c-9949-076661df0e62 · outbound

This paper cites Distributed deep learning model for intelligent video surveillance systems with edge computing.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Distributed deep learning model for intelligent video surveillance systems with edge computing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:02.499259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:02.499259Z digest=sha256:9ca356d8f13d4c458f5289176231105adc12b4169e4c8c1fff1c09e320901c0d

Observation aafc111a-b9fa-4326-966c-e495ab223121 · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Videollm-online: Online video large language model for streaming video

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:02.595350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:02.595350Z digest=sha256:f8eb1359bd7655fd0003f0a099737409caf156175ce11ee2de6e80ccdd7a434c

Observation a3ad585c-95d6-4593-ae46-fc2760f3b7ac · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Sharegpt4video: Improving video understanding and generation with better captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:02.735952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:02.735952Z digest=sha256:604ac099ea5fe5fb603027e9a8b9cc315f1b686c4835847d783abd858bd5f737

Observation 7f9f6b6c-3478-4537-9fe5-05fe9a3a0293 · outbound

This paper cites Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.720770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:02.844901Z digest=sha256:d5f52c5dea0a7a2ff94c9aa216146def0bbaa856cfe3a7486c691382058dfe35

Observation a8c5c172-dc6c-4e12-afaf-1b283327193c · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.023513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.023513Z digest=sha256:98b71e5b97ed4e9dee521d5b228ca21462122db87d22ecf6335190fe8b0ecea4

Observation 0fcbabcd-d5a3-40ab-b2f5-a01ba03c21a5 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.704419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:03.104975Z digest=sha256:33ca838686b824a20b5445f1e948ee2f54d0a8da5d02523cc793ac59026faf46

Observation 522ef2db-b67c-4aa7-924f-78c744c54477 · outbound

This paper cites Flashattention-2: Faster attention with better paral- lelism and work partitioning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Flashattention-2: Faster attention with better paral- lelism and work partitioning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.689071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:03.209972Z digest=sha256:d9ade2ba3f8ee9027132ab3a1fbef8ebe4cf1ed2db82b5993bbf7375b9e8ed11

Observation f8ab067d-b2d0-4b68-ace1-74339480aca4 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams An image is worth 16x16 words: Transformers for image recognition at scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.307705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.307705Z digest=sha256:de662db9c9db6400d5d883ac33d1d92f4753c1b4e7dca65590fb05a1dbe7edb5

Observation 2f202760-4744-4c6a-aa31-8f2171aa0c05 · outbound

This paper cites The Llama 3 Herd of Models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.374979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.374979Z digest=sha256:fd01615596ade957b6ab38842e0774e6ca185e9cba5f81f4e566d2d68becfd1e

Observation 53b89593-713c-4fb7-b545-9c19be088905 · outbound

This paper cites A density-based algorithm for discovering clusters in large spatial databases with noise.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams A density-based algorithm for discovering clusters in large spatial databases with noise

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.662907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:03.468755Z digest=sha256:aa5737ceb1448926a480f94ad39020eab73f965bb14f804dfc974db6b522d682

Observation f0d1da68-38ef-429f-8509-457ba7bebf19 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.648958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:03.534615Z digest=sha256:011a4503c0e214b5b0a2f4e0f95752b84eb8250c7c259e969d18c2c9f55df21d

Observation 726d63d5-f4ff-4c22-bf44-8fd0f11dc009 · outbound

This paper cites Temporal sentence grounding in streaming videos.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Temporal sentence grounding in streaming videos

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.635328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:03.687221Z digest=sha256:2bd90b9a96f45800de34199269768e162b03b2acd01be4e5a3c7d2ff14ef1e3b

Observation 49749f40-590f-4bee-9800-da523fe23e36 · outbound

This paper cites Mist: Multi-modal iterative spatial- temporal transformer for long-form video question answering.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Mist: Multi-modal iterative spatial- temporal transformer for long-form video question answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.621253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:03.796093Z digest=sha256:1a0b1606fae1703404554a0a86e1a1fb88770a5dbc1053b416c97ab42925db6e

Observation f38e2cb4-da5b-433f-92b1-b18c187df245 · outbound

This paper cites Clip- adapter: Better vision-language models with feature adapters.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Clip- adapter: Better vision-language models with feature adapters

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.607482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:03.874150Z digest=sha256:edd3840b6c4c3052ce4301210f08fbcec5cf86b643f1dce5a437df12accd3bd1

Observation a7e01be7-e4e9-49f0-822a-ac95500b5f06 · outbound

This paper cites Frameexit: Conditional early exiting for efficient video recognition.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Frameexit: Conditional early exiting for efficient video recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.593252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:03.969979Z digest=sha256:a91662996a4c8126a97218ea9eaab5996a690a87b55d3e60f489b24afbe35b8a

Observation 9b2576bf-59fc-4109-8f36-de56f08fe307 · outbound

This paper cites Dynamic neural networks: A survey.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Dynamic neural networks: A survey

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.580297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:04.043814Z digest=sha256:15b5fa8f901ee239d1a56b17a33f4b0f5fae766c4c1c4fd9e2409f7f631f05bc

Observation 908cd346-dabb-4b80-93c1-c82eb7e70c04 · outbound

This paper cites A twofold siamese network for real-time object tracking.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams A twofold siamese network for real-time object tracking

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.566930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:04.091875Z digest=sha256:41ce0cb4932edee2ee5a89f2b7123835e0710973a440c0f14521d9c66cc35dfc

Observation e0f959ed-ecc4-46cc-b449-3d955ef8b133 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LoRA: Low-rank adaptation of large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.554209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:04.202477Z digest=sha256:6a17f4d1dd231c26715ef157c3ca7c07aadd15e047ae5ca420eee380618ae7f2

Observation c2a9f090-f48f-4e98-b09c-0c5cd7f447a8 · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Movienet: A holistic dataset for movie understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.541398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:04.321669Z digest=sha256:ab5c9fa96f2bbfdf6509d907d5c4c4b9cd4b56cba59dec1e3b08c1b6ef143d3e

Observation 461b260e-ed3d-44a6-9ecc-b9003b49afea · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video under- standing.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Chat-univi: Unified visual representation em- powers large language models with image and video under- standing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.527716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:04.437980Z digest=sha256:b1ec0c94898117ee1a9b99f183cf1f6d25ee9bd48a5919979a00a0d3e400d0a4

Observation 20d5ba49-83be-4a5b-9471-84ffaab9a37b · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Seed-bench: Benchmarking multimodal large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.512597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:04.531283Z digest=sha256:27c6d8cbe9849d2e1f3df823dbc956a7dbf3aa2dc424d8961208b5a280791862

Observation 2bb99f58-390d-4cf3-a8b6-fd826d376847 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:04.625398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:04.625398Z digest=sha256:5b911164024337b79fb9bd1fd9f683c76f44ed364d4db895db7362697aed8e80

Observation c647090a-7381-4c06-bed0-c3a1c5d8f60c · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.498813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:04.695915Z digest=sha256:9ebf86c7961eeebefd49e9d3544ee0d74506d23238bf052b6c2e4425b2a67e64

Observation 391d44d2-f142-4fd4-9d13-ea37546e6ca8 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.484102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:04.769774Z digest=sha256:fb65fc3ce8c650f78c843fa97db4c805232d9d0fb526fd68aef8ca8f00dfdae9

Observation d3073307-a94a-44a6-b442-c23c61166907 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.470259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:04.874604Z digest=sha256:285e60a32595b3a52d21be31a395c88acaeb521874ffb88b4cb0ddb42330e3d0

Observation 82c186d8-45f2-428d-b58a-35d1f51e6d15 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Llama-vid: An image is worth 2 tokens in large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.456649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:05.041607Z digest=sha256:fef505d0384fe0bc79b07a0265d40b1f3568855c15b396ec2b563078f2eb293c

Observation 17555231-f16f-4331-b585-71124f0408f2 · outbound

This paper cites Visual instruction tuning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Visual instruction tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.444108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:05.162198Z digest=sha256:3ed88c28b35e43982994fe783dc9b3ad5dff7d6d02eb68070a506b5019c65401

Observation 1d17eb93-f65c-41fa-9216-1e97d9c997c0 · outbound

This paper cites Improved baselines with visual instruction tuning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Improved baselines with visual instruction tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.430583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:05.248330Z digest=sha256:0d6f9d52c61c9a935f10cb5f0d5c8b91961275e1f388cfcd02166b11d774f208

Observation 0f4f6773-ee34-41bd-bd60-e24713ee2a90 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:05.338028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:05.338028Z digest=sha256:20593af0a2220211a167227b5522e4ff1e3ed8b8b1ce77fc347496422b5d84f6

Observation fe3eef6e-9ba6-4747-9420-16d2f1d9c9b1 · outbound

This paper cites Learning quality-aware dynamic mem- ory for video object segmentation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Learning quality-aware dynamic mem- ory for video object segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.415628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:05.434941Z digest=sha256:88e2dc6e223248ffb024053012fa41914e1d9044268af7672ef43f7660d64b8c

Observation 181b243a-8ba6-427b-a710-d3b95a117f09 · outbound

This paper cites Universal segmentation at arbitrary granularity with language instruction.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Universal segmentation at arbitrary granularity with language instruction

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.401854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:05.607752Z digest=sha256:7329f763442ab517dee4215bc335ee4e813f716bfeef55290dff50cc79ba2b02

Observation 9a7864cc-bbaa-41c0-bd7b-dca40e139fa4 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:05.734732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:05.734732Z digest=sha256:85a629bd4811d8cefb6d76159645fbcf4637d251adcabaa40efcd1898ed656a2

Observation cc6f9ea1-d404-4abd-b1e3-0c16b116cd10 · outbound

This paper cites Soc: Semantic- assisted object cluster for referring video object segmentation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Soc: Semantic- assisted object cluster for referring video object segmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.388612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:05.827409Z digest=sha256:3dfb64f0b609a89368a2a8bd02d64e437e67054b86296ab84561e865999fe3ce

Observation 8f68580d-5ba6-40b9-a3d8-7caff45da478 · outbound

This paper cites Multi-task deep learning for real-time 3d human pose estimation and action recognition.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Multi-task deep learning for real-time 3d human pose estimation and action recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.375454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:05.934945Z digest=sha256:d60557725a59e16e3bda8589e9abd63bed4e2c38f6eabd3b018b896ab351aed4

Observation 22555c9a-deb7-4a70-8ce1-2f4a2a885b9d · outbound

This paper cites Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.362008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:06.023144Z digest=sha256:5ad693924619b9816a3ef9f3bf0bff3ade8c3b3469b3cc2af7fd86d7549fc1de

Observation be611fb7-0e21-471b-8211-369fb2e96bde · outbound

This paper cites Video-chatgpt: Towards detailed video under- standing via large vision and language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Video-chatgpt: Towards detailed video under- standing via large vision and language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.348389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:06.087437Z digest=sha256:f04749adda0ed9afb1d8b561f41e8e83345703bab35534e35a236909898c9d33

Observation 41fb913b-96b3-4fec-9d33-77a80d89d812 · outbound

This paper cites Some methods for classification and anal- ysis of multivariate observations.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Some methods for classification and anal- ysis of multivariate observations

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.334609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:06.189654Z digest=sha256:a6bba85691c1fd59bc49b0dedfa25ef8c990cd4d3e27d48be411392f03984224

Observation 7f33954e-d803-4a93-a719-e53fc98c3706 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.320839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:06.280529Z digest=sha256:ab53713d67cef53fa2444eaa3dd9e7921dea5a3714bf552ab01582ac3c43d985

Observation ba0da123-c123-4aa5-b457-02ca53d8da66 · outbound

This paper cites Deepres: A deep learning-based video summarization strategy for resource- constrained industrial surveillance scenarios.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Deepres: A deep learning-based video summarization strategy for resource- constrained industrial surveillance scenarios

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.306067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:06.351008Z digest=sha256:18a44cff422ab0ab7dcc3dd73a47a39fdc4ab97e5bdee6dfe1dd176e824d8ec5

Observation c70825d6-5f2e-4b0d-948d-316be74c920a · outbound

This paper cites Training language models to follow instructions with human feedback.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Training language models to follow instructions with human feedback

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.291436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:06.424661Z digest=sha256:7c7bab41fb7fd45816d4132a3eef1ffbea6affa9c91240d1db96d644f2f17d1d

Observation a26de3d5-160b-42bd-b8b9-9b830f97177c · outbound

This paper cites Streaming long video understanding with large language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Streaming long video understanding with large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.276943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:06.479104Z digest=sha256:7a6fe9ded80dbf24d436d98b84678646c1ad740c598edf11ac01fdf90ce21571

Observation 4b17bef0-607e-4494-8825-a72a1012c341 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.263120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:06.539428Z digest=sha256:fb4553ea83ec73b44169e7e31f0c25823ee65d8ecb1018a4ed410464cec51ff1

Observation f2a3b187-e579-4d79-abc2-17de6c4e4e55 · outbound

This paper cites Robovqa: Multimodal long-horizon reasoning for robotics.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Robovqa: Multimodal long-horizon reasoning for robotics

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.248033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:06.601443Z digest=sha256:0df1ac825f8d65f1f2780e5eaeb4bb30ab935677f9fc5c25fa73a1bda60ac47a

Observation 964afe38-b9f3-4395-b522-4bdd7d6198ae · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Moviechat: From dense token to sparse memory for long video understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.232186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:06.651780Z digest=sha256:d5a6a55b0d265e60f879844c828de4d8d78d17f084d77138510f77890e23a14a

Observation 53fc9ae4-2eb8-4020-acdd-603aa9cd8429 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Roformer: Enhanced transformer with rotary position embedding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:06.703267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:06.703267Z digest=sha256:288cff09d094e651ff1c9bf13c1a0c759556950af8b7e38729af7872161df732

Observation e2563ac0-4c65-4138-859c-d37055ffaebf · outbound

This paper cites Tracking as online decision-making: Learning a policy from streaming videos with reinforcement learning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Tracking as online decision-making: Learning a policy from streaming videos with reinforcement learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.202961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:06.822196Z digest=sha256:caf4e9b21931b6190c05fa00a0e421a00b71263de761d1a02381df52be2e56d7

Observation cfcecbf7-a784-46a9-b8e0-7978fc80d6fd · outbound

This paper cites Dynamic memory based attention network for sequential recommendation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Dynamic memory based attention network for sequential recommendation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.180383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:06.916665Z digest=sha256:9efbded5204a9424f210924f028a1f0ac55a510a5d93d0bd7c7cc922e8377ff5

Observation 6125cca0-ee24-438a-a6f3-d2b3b76defe7 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.008868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.008868Z digest=sha256:2fdad985c291d518f017b7a3f73f0a80cf730b90e21e50f1673a304ddce0d7ed

Observation b37c7254-c536-42ff-82bb-7afa328b9445 · outbound

This paper cites Kimi-VL Technical Report.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Kimi-VL Technical Report

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.074611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.074611Z digest=sha256:09c5731552182063b1c052fbbe246921e38f2260fc92edee157265855dc06b32

Observation cb49f620-fa9a-4472-9dfe-870c0cb1ac04 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LLaMA: Open and Efficient Foundation Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.127746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.127746Z digest=sha256:a79137316bc89039e286d9c8753b595b0aed111789687651ef844d025b2985b1

Observation 87ecba3e-14eb-4e4e-a8ff-8721d5970736 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.188899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.188899Z digest=sha256:95591efe86d75b6dc9c3691e57c1c5d45855a476787ed61a3da7821f61de1b46

Observation a350224f-da1f-4b50-90cf-84aaac305b3e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.239662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.239662Z digest=sha256:b008e92312d3d58584e07c67ef02c302cf84da50d0ca9a515defb9ccdfb5f58b

Observation 77574e63-3d4e-4b9b-adf3-d410f4e6881a · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LVBench: An Extreme Long Video Understanding Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.288161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.288161Z digest=sha256:92b2415c94e5a20a26fac06559b8aa28e512f73845dea56a9514dfadf7024442

Observation 242fa570-3d84-43c4-8f27-c01f37a67a4e · outbound

This paper cites ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.383398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.383398Z digest=sha256:968f7010ce9ae76a17fd358f7c9cdc431e4f789314b0304b25d6d84cae9a606a

Observation 5f9bbaf5-76ae-4e83-8855-cd7253172285 · outbound

This paper cites Adaptive focus for efficient video recognition.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Adaptive focus for efficient video recognition

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.162546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:07.498070Z digest=sha256:d854314b6545b7ce3cf7bd4c68460eb50e71b2286e01b859763037a7742574de

Observation f9a7a55a-0646-40da-a775-58ee96e5faa1 · outbound

This paper cites Adafocus v2: End-to-end training of spatial dynamic net- works for video recognition.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Adafocus v2: End-to-end training of spatial dynamic net- works for video recognition

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.146070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:07.576870Z digest=sha256:1403e8f9b8630311d152f775d365b4f1b3c32143364354d647adaf0d1007adcb

Observation 81c0ee5e-a38e-419d-a8b0-f15e86e45ecf · outbound

This paper cites Adafocusv3: On unified spatial-temporal dynamic video recognition.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Adafocusv3: On unified spatial-temporal dynamic video recognition

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.124808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:07.596032Z digest=sha256:cc04d0844a9a1fa102c2605cc353a1001b5f56a5ddb540a689e90ad8561301c5

Observation bbdda88f-2070-44df-9029-f53700619c6f · outbound

This paper cites Hierarchical Memory for Long Video QA.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Hierarchical Memory for Long Video QA

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.699169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.699169Z digest=sha256:4f65343894428ef2745825e8d7f98023a9689d8cbdc5e3b33e376c1c0e441297

Observation 7c31022a-e811-44a1-adfe-b975d579c6d7 · outbound

This paper cites Ponder & Press: Advancing Visual GUI Agent towards General Computer Control.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Ponder & Press: Advancing Visual GUI Agent towards General Computer Control

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.848372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.848372Z digest=sha256:0087d42366c18ecfd579f1262a3cf176fec50b5dbc037b4e5b5bac666ec96b3b

Observation 983783b2-0c3d-4687-88f8-c2bab012ce25 · outbound

This paper cites Uni-adafocus: Spatial-temporal dynamic computation for video recognition.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Uni-adafocus: Spatial-temporal dynamic computation for video recognition

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.110255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:08.015864Z digest=sha256:5091fb8ad9c216d75193cc47787889669e2dcee44f34aaecb6330b4fe5c66c2e

Observation 05e41abc-eb7f-4d65-90f7-facf52cbcefb · outbound

This paper cites Iterprime: Zero-shot referring image segmentation with iterative grad-cam refinement and primary word emphasis.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Iterprime: Zero-shot referring image segmentation with iterative grad-cam refinement and primary word emphasis

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.096179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:08.185558Z digest=sha256:ae131da8ff00bbd19bdd28312c4924e1120a42b15be1d33fc02ee14942a46042

Observation 108b516f-0500-4305-a3ed-f0802abfa1fa · outbound

This paper cites Sam2-love: Segment anything model 2 in language- aided audio-visual scenes.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Sam2-love: Segment anything model 2 in language- aided audio-visual scenes

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.081701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:08.352293Z digest=sha256:7e444e42aa09d6bbdd53f8d30040c72a16cf57202d0bc09cda2007ae29e6c737

Observation 9f00587a-b7bb-4c24-812d-f4a16f01545a · outbound

This paper cites Towards real-time multi-object tracking.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Towards real-time multi-object tracking

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.066009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:08.520067Z digest=sha256:de6a4ed8787b03002e21797311558e49bf20892e00a088cba8d6f82b4ac3fb1b

Observation 66990da4-9ee6-4cdd-894d-9844412dd7d1 · outbound

This paper cites Videollm-mod: Efficient video- language streaming with mixture-of-depths vision compu- tation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Videollm-mod: Efficient video- language streaming with mixture-of-depths vision compu- tation

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.048899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:08.686159Z digest=sha256:df6c0171b4edc493dfe2a28d0d3f7f3a91b5fad9f7cb6b0656b7e71e61ab517a

Observation f474a2e7-700c-40c7-beab-a09b604054f4 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Next-qa: Next phase of question-answering to explaining temporal actions

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.029636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:08.802065Z digest=sha256:2629ff36d93f379305f4c5f8677291f86d2f11720f3ae83ddba1fc50f0cc37d3

Observation 93c29c6b-51ce-4a07-9a67-99f57cf17501 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.873755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.873755Z digest=sha256:ef3c08a9e769b55d4df35e070a4d0541607ce7d0ad0eb7b61724db0e332bedb2

Observation 9eb08b61-673f-44b1-8016-d638a7fe5d3d · outbound

This paper cites Fine-grained video captioning via graph-based multi- granularity interaction learning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Fine-grained video captioning via graph-based multi- granularity interaction learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.012337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:08.879930Z digest=sha256:8a266d839cb2d289aa199dedf745f70eb94b9b0b9502a3249f5a74a42937b6d2

Observation b050451d-f53c-491f-950b-74967d66ffce · outbound

This paper cites Qwen2 Technical Report.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Qwen2 Technical Report

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.890274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.890274Z digest=sha256:57a208a2f1ab99126978c6175176158899821246fb25efa87e92dd8041d9578d

Observation fbe6b9ba-319a-4d92-9783-f9ee5860840b · outbound

This paper cites Language-aware vision transformer for referring segmentation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Language-aware vision transformer for referring segmentation

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.994052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:08.900750Z digest=sha256:3f50b53028e202d7108c5da3ad3f15935b399da0f0fc3d1f836c99f215aafbd9

Observation 1cf98fa4-906c-44af-af7c-e0594050c865 · outbound

This paper cites Atp-llava: Adaptive token pruning for large vision language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Atp-llava: Adaptive token pruning for large vision language models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.975487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:08.917306Z digest=sha256:e3f9f504a88778d7830f185bdfd1a2dd7d1f741c446ee689a3d0d7bf7a048ce9

Observation d93939c1-ba4a-49c8-bdbf-69ed035ef2a7 · outbound

This paper cites V oco-llama: Towards vision compression with large language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams V oco-llama: Towards vision compression with large language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.958375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:08.923634Z digest=sha256:5bee25da38533eb5b0e0928a2fc59648d24e52be921bcf8f0ac0ee77aea36bd4

Observation f40af02a-f349-43ba-b593-e05875441c11 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Self-chained image-language model for video localization and question answering

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.933447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:08.934527Z digest=sha256:6a2c410aa41ac9ae872f4169193055c7969b767ec8044b07bb4eedb875854ad3

Observation 5708aa3f-d747-4b72-9ae2-314e41a9b0a1 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.913146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:08.942219Z digest=sha256:353bd996ffb13c39d948fe68d0fd07ad6adfb2c42f4f4c9701f9c5c99a808e06

Observation 3b095665-9420-4c13-8203-70c5b0f8f2b4 · outbound

This paper cites Real-time action recognition with enhanced motion vector cnns.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Real-time action recognition with enhanced motion vector cnns

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.892300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:08.956859Z digest=sha256:468c2660a20d474e99f083384f8c4b272c4b31cfe7ee344605a5c98c95a740b4

Observation fafdc095-d006-47ba-8474-889537b0e4ff · outbound

This paper cites Long Context Transfer from Language to Vision.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Long Context Transfer from Language to Vision

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.962828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.962828Z digest=sha256:7bcccfb5dd2a31a7ba01abce1bd2576387657b07df49b5a6bde6ba5559fab682

Observation 113072f7-b149-48b4-98a8-8e9f4e1254e5 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.967515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.967515Z digest=sha256:31bc672253daa2549878801602877bb0bf9410c3abc73433e4159c8c09eee183

Observation 52e61739-f8b5-43b7-a14c-62e6067bddfc · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams MLVU: Benchmarking Multi-task Long Video Understanding

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.972286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.972286Z digest=sha256:5f85de197d8d704eece507c3a22cba7c69cebfea1e98c1da1d2b3996d45eb628

Observation 2c8153d6-b46f-4322-b8f9-89626a8f23b0 · outbound

This paper cites Streaming dense video captioning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Streaming dense video captioning

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.875831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:08.977508Z digest=sha256:e5d3aed27aadc16c40e4bdd61d12e47b39b594f938236b1f878837d6cb73730a

Observation def39505-06cc-4ccf-baed-4dff0a300a59 · outbound

This paper cites InstaRevive: One-Step Image Enhancement via Dynamic Score Matching.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams InstaRevive: One-Step Image Enhancement via Dynamic Score Matching

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:37:09.053145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:08.984190Z digest=sha256:c12c4ab06dc6cfc1d0615f7d193932e142429bf162c289a6c44576a15ef07dda

Observation e10025c3-8e4e-4ce7-8f23-a47b69e63cce · outbound

This paper cites an unresolved cited work.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Unresolved cited work

Reference 1000

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:37:09.856397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:37:08.992465Z digest=sha256:5e36f595f292d2513615ae3399d5273bd63ecb8ac0219754aa2b2031e93dff24

Pith citing papers

Observation 5d958aba-4e18-443b-bce7-01c2b8959940 · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:35.208400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:35.208400Z digest=sha256:008353b57d835877f411e846389027d28e2068e317f3ae86f3984e69a6573d7c

Observation c89077fc-d844-4efe-80f0-208470118693 · inbound

ScoreHOI: Physically Plausible Reconstruction of Human-Object Interaction via Score-Guided Diffusion cites this paper.

ScoreHOI: Physically Plausible Reconstruction of Human-Object Interaction via Score-Guided Diffusion Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T16:14:01.995575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:14:01.995575Z digest=sha256:3fa8f2bcc825296fb5f264da8d88a4a6707eed4b5ea4ea87107b7fb9ee5db9ca

Observation 011ffc0d-af10-4df5-b428-909d0d7ea02e · inbound

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation cites this paper.

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T14:32:00.456065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:32:00.456065Z digest=sha256:f1f7b8da876dcda302c46e498683da8e173a3b5f686b5cbd7f846aa4b880e41e

Observation 6f3dff6f-4043-41f6-b592-79e133eaebbc · inbound

Mosaic: Cross-Modal Clustering for Efficient Video Understanding cites this paper.

Mosaic: Cross-Modal Clustering for Efficient Video Understanding Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T16:10:34.388472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:07:36.133404Z digest=sha256:53e533ad802159ab7cf139761d48171f7f58feb6e56f6d95eb1811b802a97c59

Observation ce12786c-d594-42aa-8810-0d56971e0a6c · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:03.709919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:bf78c9a4be52faaec6fb120ee53caa86fa4ee2cfcb28e499835b8d46dea54b01

Observation 78041cb7-9628-4162-a2e3-c47795275266 · inbound

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning cites this paper.

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.780231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T06:51:52.861981Z digest=sha256:0beff38374051723c4f354d5374551bd6de7060e19e0a62a6b03ce16a10b0aeb

Observation 44ceb307-4b36-4750-bc72-23d3b53ade81 · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:20:56.393925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T02:30:55.939351Z digest=sha256:a2dbb9e22c8fdc8798c81aa0729438c23b7c258049b80c9c180b420eba793f35

Observation cf4b5ff5-efc3-4be8-997b-d36a18201cae · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:01:17.733473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:00be1847366e88b8e8dfe2426a616700b2a79c0f9de03e74a540d2a2a68226c3

Observation 54e3e141-b2f2-4945-be55-c0e5ffa8023f · inbound

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering cites this paper.

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-06-28T02:01:29.198189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T01:52:33.768494Z digest=sha256:9412e182a294efc4d1f774b548e584444c1303699d859e83d1d8234a28a424bb

Observation 061c197b-c067-4d9a-bdf5-0e5971b2fd3a · inbound

Harnessing Streaming Video in the Wild cites this paper.

Harnessing Streaming Video in the Wild Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:37:25.733342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T18:47:55.910417Z digest=sha256:55f9739c7eb3884ad5480cd2471c79e01ed85e78f6287055fa08b1711149e35c

Observation e4615993-e9ff-4785-81b1-0e0ca671466b · inbound

Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation cites this paper.

Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-14T04:39:27.358904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:39:27.358904Z digest=sha256:7abc919a8f3b84b3698e774fdd2ccb48d5775512576bed2aae1c0f1b582b0d14