Pith. sign in

Paper Citation Record · LEDGER

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

As of 8 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 9 inbound Pith citation observations for arXiv:2506.23825.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23825 v2

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:37:08.992465Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:03:35.208400Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact2
  • verified fuzzy54
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6887147d-f18e-4ce0-be0b-649e9ee97769 · outbound

This paper cites GPT-4 Technical Report.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:01.869344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:01.869344Z digest=sha256:b58c4302949e28fd3cf074c459906a110be81b90879259aa7a8a3e2d9866b99a

Observation f951860e-258b-40dd-9b3e-695ac42e79ff · outbound

This paper cites Self-calibrated clip for training-free open-vocabulary segmentation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Self-calibrated clip for training-free open-vocabulary segmentation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:01.948138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:01.948138Z digest=sha256:ecbd247566c97824a6c43e10215cc5271d7e538c7e87b7922c16f170537f6505

Observation e8c54936-2933-4df9-a915-0306862a540a · outbound

This paper cites Memory consolidation enables long-context video understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Memory consolidation enables long-context video understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:02.070349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:02.070349Z digest=sha256:c039a83e0b57c4ccb31e18c750403c095ee4dbbe2dcb068f8921e89304cb4ca3

Observation f1144657-4b06-4101-b70e-f10325e7f806 · outbound

This paper cites Language models are few-shot learners.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Language models are few-shot learners

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:02.199304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:02.199304Z digest=sha256:fae9e21aa4c5c80a1690225fce77aef7c4d249f8da479decaa4b7f5efe5c6c8e

Observation ed047457-67d2-493f-b1bd-4a5a5b54c0da · outbound

This paper cites A Memory-Network Based Solution for Multivariate Time-Series Forecasting.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams A Memory-Network Based Solution for Multivariate Time-Series Forecasting

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:37:09.602810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:02.378276Z digest=sha256:c7437944cbc252d1ddea6a3491321df5505894f772430e6983c71e3fba98785d

Observation caa6d757-081e-424c-9949-076661df0e62 · outbound

This paper cites Distributed deep learning model for intelligent video surveillance systems with edge computing.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Distributed deep learning model for intelligent video surveillance systems with edge computing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:02.499259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:02.499259Z digest=sha256:3593d7b4f88a32d80ac741bae6824f297a2dbd86bb35d7f3765d806e29e28455

Observation aafc111a-b9fa-4326-966c-e495ab223121 · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Videollm-online: Online video large language model for streaming video

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:02.595350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:02.595350Z digest=sha256:124381c14900cac63cf3bcf4d2fc3ecb168ff91233db69eb8ebb4623da0f289f

Observation a3ad585c-95d6-4593-ae46-fc2760f3b7ac · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Sharegpt4video: Improving video understanding and generation with better captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:02.735952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:02.735952Z digest=sha256:ba9d4e64bd6f64605763baae65f65bfdb8760c1adf2ec3b2626f1a417772de2a

Observation 7f9f6b6c-3478-4537-9fe5-05fe9a3a0293 · outbound

This paper cites Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.720770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:02.844901Z digest=sha256:ca9b5c35a40b780de1febe3362a8840ac601f529e5d113dfb061bf75d28ea594

Observation a8c5c172-dc6c-4e12-afaf-1b283327193c · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.023513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.023513Z digest=sha256:5dbccb0d64a1d2bc738422aa90a4fa3652f816dfceff14fc41ebae945275eb4a

Observation 0fcbabcd-d5a3-40ab-b2f5-a01ba03c21a5 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.704419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:03.104975Z digest=sha256:aeaf22e9bbb508f23f9492722b32c544f27cfd9496cc5adb3e7411e48e1e405c

Observation 522ef2db-b67c-4aa7-924f-78c744c54477 · outbound

This paper cites Flashattention-2: Faster attention with better paral- lelism and work partitioning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Flashattention-2: Faster attention with better paral- lelism and work partitioning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.689071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:03.209972Z digest=sha256:e92c780457800fce04cee77059074a28187c96b98eecde3925ddcd6417668040

Observation f8ab067d-b2d0-4b68-ace1-74339480aca4 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams An image is worth 16x16 words: Transformers for image recognition at scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.307705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.307705Z digest=sha256:89fc2ce8c1484ba5620d4c10959a82ff8ca9cb8fd66be8c5c59e060cdab01aa5

Observation 2f202760-4744-4c6a-aa31-8f2171aa0c05 · outbound

This paper cites The Llama 3 Herd of Models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.374979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.374979Z digest=sha256:4b1e35aa0ac642e9b853f31d593ae7c8b8d47d4cd14aaa6024e6bbe01b3a98f7

Observation 53b89593-713c-4fb7-b545-9c19be088905 · outbound

This paper cites A density-based algorithm for discovering clusters in large spatial databases with noise.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams A density-based algorithm for discovering clusters in large spatial databases with noise

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.662907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:03.468755Z digest=sha256:0f986afbba9751fb0dc72675a3eeb383a8f1103a9c5c2769a812cb222d128791

Observation f0d1da68-38ef-429f-8509-457ba7bebf19 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.648958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:03.534615Z digest=sha256:f841a737981fac9b73c26646339f2aabaac16c906e5fd35ab00e3c9cdbff1454

Observation 726d63d5-f4ff-4c22-bf44-8fd0f11dc009 · outbound

This paper cites Temporal sentence grounding in streaming videos.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Temporal sentence grounding in streaming videos

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.635328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:03.687221Z digest=sha256:c6b3015c10acf3839d1f383139ea777ff982dd9f02a084a76a87f750ccea29cb

Observation 49749f40-590f-4bee-9800-da523fe23e36 · outbound

This paper cites Mist: Multi-modal iterative spatial- temporal transformer for long-form video question answering.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Mist: Multi-modal iterative spatial- temporal transformer for long-form video question answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.621253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:03.796093Z digest=sha256:45f393b21409d64f2023cce28a6986139b778206e3a33eee05b81aa7842e5001

Observation f38e2cb4-da5b-433f-92b1-b18c187df245 · outbound

This paper cites Clip- adapter: Better vision-language models with feature adapters.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Clip- adapter: Better vision-language models with feature adapters

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.607482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:03.874150Z digest=sha256:f751464e3a7cfa145fcf8d842fae0ea7583810875bddd17178c29a43f57a70a7

Observation a7e01be7-e4e9-49f0-822a-ac95500b5f06 · outbound

This paper cites Frameexit: Conditional early exiting for efficient video recognition.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Frameexit: Conditional early exiting for efficient video recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.593252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:03.969979Z digest=sha256:5bee70536058c8f6f6439d54bddd52814ee31b4aec2cff49e8edf06cc9d1cb45

Observation 9b2576bf-59fc-4109-8f36-de56f08fe307 · outbound

This paper cites Dynamic neural networks: A survey.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Dynamic neural networks: A survey

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.580297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:04.043814Z digest=sha256:f360750ca1e9f0727a47e5037cc57961eca475968707ff867dfff635ee6f5005

Observation 908cd346-dabb-4b80-93c1-c82eb7e70c04 · outbound

This paper cites A twofold siamese network for real-time object tracking.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams A twofold siamese network for real-time object tracking

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.566930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:04.091875Z digest=sha256:7307c12552e5107cdfdbdf5a93762f0ebb95b728bcf4c742e506e028672b6bae

Observation e0f959ed-ecc4-46cc-b449-3d955ef8b133 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LoRA: Low-rank adaptation of large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.554209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:04.202477Z digest=sha256:8d065726cf25e70535890febcd8ef700421c3a0aa3c2e80eb6e754ae1f325942

Observation c2a9f090-f48f-4e98-b09c-0c5cd7f447a8 · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Movienet: A holistic dataset for movie understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.541398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:04.321669Z digest=sha256:f14af4ec18add94dce9e6109cb6640811a87ab611815e48b35fcbfa146b650e7

Observation 461b260e-ed3d-44a6-9ecc-b9003b49afea · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video under- standing.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Chat-univi: Unified visual representation em- powers large language models with image and video under- standing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.527716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:04.437980Z digest=sha256:7b192111493e136874784f1478c08621718e15c32c953564002d31b30912a19a

Observation 20d5ba49-83be-4a5b-9471-84ffaab9a37b · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Seed-bench: Benchmarking multimodal large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.512597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:04.531283Z digest=sha256:8d1e6711ee61c833035e618e1563a5b31a73084aaf182037fc4c837c485e627f

Observation 2bb99f58-390d-4cf3-a8b6-fd826d376847 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:04.625398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:04.625398Z digest=sha256:e6def4eacbdeec75add99c6cb73101375d79445d5ab896ab6a728d1f2fdaf90a

Observation c647090a-7381-4c06-bed0-c3a1c5d8f60c · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.498813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:04.695915Z digest=sha256:0acaaa9c1ac364ffc0f0b9078366192becd6650f030f46418f755e0a004b6988

Observation 391d44d2-f142-4fd4-9d13-ea37546e6ca8 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.484102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:04.769774Z digest=sha256:717dc6f4feddb00fc1b7a34f97c3f2d2ff567134f4ba46a13a8f9b946abb29c1

Observation d3073307-a94a-44a6-b442-c23c61166907 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.470259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:04.874604Z digest=sha256:8d46202b55d3f3bb4598daeff1d93964ec413c3d047ec8a24a8fb28300b998a2

Observation 82c186d8-45f2-428d-b58a-35d1f51e6d15 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Llama-vid: An image is worth 2 tokens in large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.456649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:05.041607Z digest=sha256:48c25897c173f9803eb611f04d5fa13da152bb94b665194e828cc744571633a6

Observation 17555231-f16f-4331-b585-71124f0408f2 · outbound

This paper cites Visual instruction tuning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Visual instruction tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.444108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:05.162198Z digest=sha256:7754d3c67a0ea84bf9bc645e46b0ee184a94d7e6616744753847dbc4e657e724

Observation 1d17eb93-f65c-41fa-9216-1e97d9c997c0 · outbound

This paper cites Improved baselines with visual instruction tuning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Improved baselines with visual instruction tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.430583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:05.248330Z digest=sha256:00fdf5db05580a32ad7eba9b21f20ad5dc04736d240b2e8eadd6d7f1bf03c687

Observation 0f4f6773-ee34-41bd-bd60-e24713ee2a90 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:05.338028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:05.338028Z digest=sha256:42a3e1e7bd47aef9c96da20f1becd53af9816aacc51dbc99bb2acbe051703e18

Observation fe3eef6e-9ba6-4747-9420-16d2f1d9c9b1 · outbound

This paper cites Learning quality-aware dynamic mem- ory for video object segmentation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Learning quality-aware dynamic mem- ory for video object segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.415628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:05.434941Z digest=sha256:bdf1fe13531bf0626bbdbc53de5cb6620f329ccb4a74cc1b8e2508677bfb3153

Observation 181b243a-8ba6-427b-a710-d3b95a117f09 · outbound

This paper cites Universal segmentation at arbitrary granularity with language instruction.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Universal segmentation at arbitrary granularity with language instruction

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.401854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:05.607752Z digest=sha256:f6e56fd1cae70a3562f3a76b7a3440429199cdd5db0bb1b6648f39f47e1c00da

Observation 9a7864cc-bbaa-41c0-bd7b-dca40e139fa4 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:05.734732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:05.734732Z digest=sha256:97208d89108e8fa8c2060e2d34b367f09559e47c63895ccfa6b290002e23ad02

Observation cc6f9ea1-d404-4abd-b1e3-0c16b116cd10 · outbound

This paper cites Soc: Semantic- assisted object cluster for referring video object segmentation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Soc: Semantic- assisted object cluster for referring video object segmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.388612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:05.827409Z digest=sha256:6ff26f63c51a5f5d46418b958f83505b873f69c24709bf6fcb24985243170586

Observation 8f68580d-5ba6-40b9-a3d8-7caff45da478 · outbound

This paper cites Multi-task deep learning for real-time 3d human pose estimation and action recognition.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Multi-task deep learning for real-time 3d human pose estimation and action recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.375454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:05.934945Z digest=sha256:79705850f613d999a59159b010af4ba150d3262d3aab5e4064e4c69c869713d7

Observation 22555c9a-deb7-4a70-8ce1-2f4a2a885b9d · outbound

This paper cites Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.362008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.023144Z digest=sha256:ad8cd461cc39a07118758478d8e3fb7315b23e5c5953c9a9e94beeedfcf5903a

Observation be611fb7-0e21-471b-8211-369fb2e96bde · outbound

This paper cites Video-chatgpt: Towards detailed video under- standing via large vision and language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Video-chatgpt: Towards detailed video under- standing via large vision and language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.348389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.087437Z digest=sha256:899a2886d4ad3a6c07d0266ca5f1124c10dba58097c5d89e50d906e0009ee8c2

Observation 41fb913b-96b3-4fec-9d33-77a80d89d812 · outbound

This paper cites Some methods for classification and anal- ysis of multivariate observations.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Some methods for classification and anal- ysis of multivariate observations

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.334609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.189654Z digest=sha256:a1013c4350b3acd09bc558f06c29888511108f81ee773c59e3d98fd22b00c1c3

Observation 7f33954e-d803-4a93-a719-e53fc98c3706 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.320839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.280529Z digest=sha256:90ce2284f0c15654def383bd8fbbcdb7d06664ee40603693baa054efa305f2f2

Observation ba0da123-c123-4aa5-b457-02ca53d8da66 · outbound

This paper cites Deepres: A deep learning-based video summarization strategy for resource- constrained industrial surveillance scenarios.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Deepres: A deep learning-based video summarization strategy for resource- constrained industrial surveillance scenarios

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.306067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.351008Z digest=sha256:cfe3a8ae1dc9356441d53644c9d20b0bb97b2d404ea090c92eef3c93d9185b65

Observation c70825d6-5f2e-4b0d-948d-316be74c920a · outbound

This paper cites Training language models to follow instructions with human feedback.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Training language models to follow instructions with human feedback

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.291436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.424661Z digest=sha256:649cebbe3dad65f69d9f5a149ebde5c183e61ff4919373b8eb8d3cc34e88719e

Observation a26de3d5-160b-42bd-b8b9-9b830f97177c · outbound

This paper cites Streaming long video understanding with large language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Streaming long video understanding with large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.276943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.479104Z digest=sha256:c652385853c5a4375259b7ddd9d6cda3d68c6c962da6b3e45d1d926f4d9453b0

Observation 4b17bef0-607e-4494-8825-a72a1012c341 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.263120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.539428Z digest=sha256:36be0bda6385c74bf5953b0b029a5f04e29b472d1ba409fd91ef266a852060e3

Observation f2a3b187-e579-4d79-abc2-17de6c4e4e55 · outbound

This paper cites Robovqa: Multimodal long-horizon reasoning for robotics.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Robovqa: Multimodal long-horizon reasoning for robotics

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.248033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.601443Z digest=sha256:b9fca6100e6a91cbe201b15179e54f9175147d5bed60a4d1ba46497bf4b7a42a

Observation 964afe38-b9f3-4395-b522-4bdd7d6198ae · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Moviechat: From dense token to sparse memory for long video understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.232186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.651780Z digest=sha256:45840aab49ffb829e427cde64eaf9fa4c153708965657132ee63eba509e5e437

Observation 53fc9ae4-2eb8-4020-acdd-603aa9cd8429 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Roformer: Enhanced transformer with rotary position embedding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:06.703267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:06.703267Z digest=sha256:a1b89c86368aa81e402d746340f0d273c509366c748427e4ad6342d0d6270f0e

Observation e2563ac0-4c65-4138-859c-d37055ffaebf · outbound

This paper cites Tracking as online decision-making: Learning a policy from streaming videos with reinforcement learning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Tracking as online decision-making: Learning a policy from streaming videos with reinforcement learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.202961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.822196Z digest=sha256:e4ce689bd1500f3a71b029a46f06e94830ea2a21d95276b81f9207c4e07dc383

Observation cfcecbf7-a784-46a9-b8e0-7978fc80d6fd · outbound

This paper cites Dynamic memory based attention network for sequential recommendation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Dynamic memory based attention network for sequential recommendation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.180383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.916665Z digest=sha256:99035045cfdfa7ac58a40c4c751255e66ea6734b2fc5713967ca5bc62c8384f7

Observation 6125cca0-ee24-438a-a6f3-d2b3b76defe7 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.008868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.008868Z digest=sha256:a617524ce1e558349dcecd97978d8f88472e4f9a6aabe93277afb5b7f2a158fd

Observation b37c7254-c536-42ff-82bb-7afa328b9445 · outbound

This paper cites Kimi-VL Technical Report.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Kimi-VL Technical Report

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.074611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.074611Z digest=sha256:57fd48bc726809b153dc7b64104ed7c95f5833adadf676bba1746ff12a2ed77d

Observation cb49f620-fa9a-4472-9dfe-870c0cb1ac04 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LLaMA: Open and Efficient Foundation Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.127746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.127746Z digest=sha256:9a216e4440fa8b73bd01042b20853a3336495943a9e1043f5161d89e0a2e0bf7

Observation 87ecba3e-14eb-4e4e-a8ff-8721d5970736 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.188899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.188899Z digest=sha256:823f2c52c809e577856e1af73fea1e2a6fe44b01a25b0bbf0fe4ed6c9eb2970e

Observation a350224f-da1f-4b50-90cf-84aaac305b3e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.239662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.239662Z digest=sha256:76e4a2a0939ee290b3960f068bba3474b4908de83038c3098c9f7be9adc41ff1

Observation 77574e63-3d4e-4b9b-adf3-d410f4e6881a · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LVBench: An Extreme Long Video Understanding Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.288161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.288161Z digest=sha256:eca556dc5d9ab0ed3217f041491e7e918638b18b3f8e718e2a855bb4068353b2

Observation 242fa570-3d84-43c4-8f27-c01f37a67a4e · outbound

This paper cites ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.383398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.383398Z digest=sha256:092c2d03a361c66b0e82a9425a1fd3bfc620e97c2fb6a42140fc0a0db2fe775b

Observation 5f9bbaf5-76ae-4e83-8855-cd7253172285 · outbound

This paper cites Adaptive focus for efficient video recognition.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Adaptive focus for efficient video recognition

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.162546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:07.498070Z digest=sha256:d53e56f68da5fab1452171708b371aaebe17733b989182533cc5b1caecada963

Observation f9a7a55a-0646-40da-a775-58ee96e5faa1 · outbound

This paper cites Adafocus v2: End-to-end training of spatial dynamic net- works for video recognition.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Adafocus v2: End-to-end training of spatial dynamic net- works for video recognition

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.146070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:07.576870Z digest=sha256:53fb49113aa416ccc217e82cec8f3ae9f063503855afd447ccf32faf6787da8d

Observation 81c0ee5e-a38e-419d-a8b0-f15e86e45ecf · outbound

This paper cites Adafocusv3: On unified spatial-temporal dynamic video recognition.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Adafocusv3: On unified spatial-temporal dynamic video recognition

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.124808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:07.596032Z digest=sha256:34b8d00a1bc37c81e1dbca1b9ef7218cceef10d855134676f3e189c89e08f2b8

Observation bbdda88f-2070-44df-9029-f53700619c6f · outbound

This paper cites Hierarchical Memory for Long Video QA.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Hierarchical Memory for Long Video QA

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.699169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.699169Z digest=sha256:1b6d3f1d926bf67667a97a6be04d2086908e82b4ba7083d4cda5ac4595535bff

Observation 7c31022a-e811-44a1-adfe-b975d579c6d7 · outbound

This paper cites Ponder & Press: Advancing Visual GUI Agent towards General Computer Control.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Ponder & Press: Advancing Visual GUI Agent towards General Computer Control

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.848372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.848372Z digest=sha256:fca3d95809ad15613879fa53760ddfdec8195c2b807c97504e558819203bed49

Observation 983783b2-0c3d-4687-88f8-c2bab012ce25 · outbound

This paper cites Uni-adafocus: Spatial-temporal dynamic computation for video recognition.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Uni-adafocus: Spatial-temporal dynamic computation for video recognition

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.110255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.015864Z digest=sha256:f4eeff2514a34fb0b4097f3ead415baed0c74d240c7e2fb403bffad215e9883e

Observation 05e41abc-eb7f-4d65-90f7-facf52cbcefb · outbound

This paper cites Iterprime: Zero-shot referring image segmentation with iterative grad-cam refinement and primary word emphasis.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Iterprime: Zero-shot referring image segmentation with iterative grad-cam refinement and primary word emphasis

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.096179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.185558Z digest=sha256:849292d73291e12eae7308475022989222af214f4a18d452643802d78e13499f

Observation 108b516f-0500-4305-a3ed-f0802abfa1fa · outbound

This paper cites Sam2-love: Segment anything model 2 in language- aided audio-visual scenes.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Sam2-love: Segment anything model 2 in language- aided audio-visual scenes

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.081701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.352293Z digest=sha256:a1e90e92f742e4bfdb0fd66e2aff7e50bfbd60e24840acf472f49d925805f82d

Observation 9f00587a-b7bb-4c24-812d-f4a16f01545a · outbound

This paper cites Towards real-time multi-object tracking.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Towards real-time multi-object tracking

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.066009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.520067Z digest=sha256:0d5e20650f31dbc0a4bdf907e4ab42f7effa3547573f858eaf058c50201197fd

Observation 66990da4-9ee6-4cdd-894d-9844412dd7d1 · outbound

This paper cites Videollm-mod: Efficient video- language streaming with mixture-of-depths vision compu- tation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Videollm-mod: Efficient video- language streaming with mixture-of-depths vision compu- tation

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.048899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.686159Z digest=sha256:92632f1b1a6fdb280301004c013078870a64036a470aaf544acb2fbfda33a707

Observation f474a2e7-700c-40c7-beab-a09b604054f4 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Next-qa: Next phase of question-answering to explaining temporal actions

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.029636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.802065Z digest=sha256:f84bccd127ccf4c76ac46142e3b163f7e47bda5977654213784f6ccd57ffca6b

Observation 93c29c6b-51ce-4a07-9a67-99f57cf17501 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.873755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.873755Z digest=sha256:bcd8a483eb7bd642d3ff0ce62880d6b91292c9033631f40a3291abcd5c5d8b28

Observation 9eb08b61-673f-44b1-8016-d638a7fe5d3d · outbound

This paper cites Fine-grained video captioning via graph-based multi- granularity interaction learning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Fine-grained video captioning via graph-based multi- granularity interaction learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.012337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.879930Z digest=sha256:4138b7ba2c6916ba8599cad507b0d131b9a26c2885de4d90f695cb03f70e532b

Observation b050451d-f53c-491f-950b-74967d66ffce · outbound

This paper cites Qwen2 Technical Report.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Qwen2 Technical Report

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.890274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.890274Z digest=sha256:c9e870f55fad081eedce1f076d05e6fe30545c1e60e0c6b960629071f6825614

Observation fbe6b9ba-319a-4d92-9783-f9ee5860840b · outbound

This paper cites Language-aware vision transformer for referring segmentation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Language-aware vision transformer for referring segmentation

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.994052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.900750Z digest=sha256:81987af7b15ce683c4821504adc2f89c92c445a14d896b29bd46080aad68daa4

Observation 1cf98fa4-906c-44af-af7c-e0594050c865 · outbound

This paper cites Atp-llava: Adaptive token pruning for large vision language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Atp-llava: Adaptive token pruning for large vision language models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.975487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.917306Z digest=sha256:f482e6fef8e4f2d2be0c30f3151d13130808180d15d81c87c70293f9e6f97eb9

Observation d93939c1-ba4a-49c8-bdbf-69ed035ef2a7 · outbound

This paper cites V oco-llama: Towards vision compression with large language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams V oco-llama: Towards vision compression with large language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.958375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.923634Z digest=sha256:5df91e48b0dfb0427db082014f0780e73dcb7094cdcc06f4a2e5d0dab4cf39cf

Observation f40af02a-f349-43ba-b593-e05875441c11 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Self-chained image-language model for video localization and question answering

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.933447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.934527Z digest=sha256:5b644f00013c611cc34baa8912f4b0956a74d5600b3afac2afd2aff189096db2

Observation 5708aa3f-d747-4b72-9ae2-314e41a9b0a1 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.913146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.942219Z digest=sha256:6c0bf7ee1afe01c853e00b22c4033e42b77d05059d0a6d3503290cba5c4d891f

Observation 3b095665-9420-4c13-8203-70c5b0f8f2b4 · outbound

This paper cites Real-time action recognition with enhanced motion vector cnns.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Real-time action recognition with enhanced motion vector cnns

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.892300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.956859Z digest=sha256:d4340ed934911e11293ee35018603d4b8cc87131074bed42b91c4a7d5e2a3e3f

Observation fafdc095-d006-47ba-8474-889537b0e4ff · outbound

This paper cites Long Context Transfer from Language to Vision.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Long Context Transfer from Language to Vision

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.962828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.962828Z digest=sha256:a3f765d87c31b7d1e410ce404d7c1b0adcb2b9fac5dd3903cf74526529f7d863

Observation 113072f7-b149-48b4-98a8-8e9f4e1254e5 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.967515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.967515Z digest=sha256:2ff32160bd091744859ca76a500e8dc20a35100af6d960886ae990cb343fe5f9

Observation 52e61739-f8b5-43b7-a14c-62e6067bddfc · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams MLVU: Benchmarking Multi-task Long Video Understanding

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.972286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.972286Z digest=sha256:253490e4a608a25c205661eed1e6d0f05d9c3a142ab03a42d08f9ed4bc46adda

Observation 2c8153d6-b46f-4322-b8f9-89626a8f23b0 · outbound

This paper cites Streaming dense video captioning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Streaming dense video captioning

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.875831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.977508Z digest=sha256:722ae22646776bc4b6642e340444b7fb93ab443bc9dd790b38177f05cf34b137

Observation def39505-06cc-4ccf-baed-4dff0a300a59 · outbound

This paper cites InstaRevive: One-Step Image Enhancement via Dynamic Score Matching.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams InstaRevive: One-Step Image Enhancement via Dynamic Score Matching

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:37:09.053145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.984190Z digest=sha256:4384f9a370b6c7c3fbf0d457148397282fa83b6598d5b8227ba9dccd705d20f9

Observation e10025c3-8e4e-4ce7-8f23-a47b69e63cce · outbound

This paper cites an unresolved cited work.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Unresolved cited work

Reference 1000

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:37:09.856397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.992465Z digest=sha256:5378023f509e99b3c2ad31d1d3bdc2e928fda498a7f7b9419c66f9a2fb226807

Pith citing papers

Observation 5d958aba-4e18-443b-bce7-01c2b8959940 · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:35.208400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:35.208400Z digest=sha256:679bd1a8c4978fe5f46f37198bfa8ead468526f8cc0e723a7268e8f14823445f

Observation 011ffc0d-af10-4df5-b428-909d0d7ea02e · inbound

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation cites this paper.

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T14:32:00.456065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:32:00.456065Z digest=sha256:39ee106d82eb712b077d54f5cf2f090ee986b1d4a49d68cbe018f39114649411

Observation 6f3dff6f-4043-41f6-b592-79e133eaebbc · inbound

Mosaic: Cross-Modal Clustering for Efficient Video Understanding cites this paper.

Mosaic: Cross-Modal Clustering for Efficient Video Understanding Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T16:10:34.388472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:07:36.133404Z digest=sha256:25c891119653094db0bff0e71de363d694c178d332a5da2a931e41c120b5420b

Observation ce12786c-d594-42aa-8810-0d56971e0a6c · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:03.709919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:7bf95f14826ec2348b7f494faa635b5369c816b925e8b51d4ff44d71e8ff522e

Observation 78041cb7-9628-4162-a2e3-c47795275266 · inbound

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning cites this paper.

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.780231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T06:51:52.861981Z digest=sha256:9d344db9a193bec7347f0695c6c95ac44ac85a87d0e57e30dace152eeb025f0f

Observation 44ceb307-4b36-4750-bc72-23d3b53ade81 · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:20:56.393925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T02:30:55.939351Z digest=sha256:ff2c70a865b7ff8afd80384f591fe6088fc7483aa0dd60d09608216f34799c46

Observation cf4b5ff5-efc3-4be8-997b-d36a18201cae · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:01:17.733473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:c752faf0b8fc9ad46b19c36668cfa46e027425777ba9229f761cdc721a3f22c8

Observation 54e3e141-b2f2-4945-be55-c0e5ffa8023f · inbound

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering cites this paper.

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-06-28T02:01:29.198189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T01:52:33.768494Z digest=sha256:f204439af09b9d2d2c6e939f1684a8ab8dff300099823ecb7825840fc805363a

Observation 061c197b-c067-4d9a-bdf5-0e5971b2fd3a · inbound

Harnessing Streaming Video in the Wild cites this paper.

Harnessing Streaming Video in the Wild Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:37:25.733342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T18:47:55.910417Z digest=sha256:5b6c32747a36b54099f37c11ae30a50d792627eb52191a77b427cc253c8796a5