Pith. sign in

Paper Citation Record · LEDGER

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos

As of 18 August 2026, this Paper Citation Record lists 97 of 97 outbound references and 9 inbound Pith citation observations for arXiv:2504.17343.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.17343 v1

Coverage vector

measured 97 of 97 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:47:27.549539Z

measured 106 of 106 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:39:27.373948Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T05:39:06.193766Z

Reference resolution

97 of 97 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved88
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24a0a1e7-953d-4007-87f5-dc7bf1ba47c3 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.087191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.087191Z digest=sha256:d12dab0c3b186df896a38f85802b2ce68ba42680dae717327007c627794a4f42

Observation 85a5fc5a-74eb-4099-984c-0899177f66e3 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.093080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.093080Z digest=sha256:a598515aed9d0826602384dbc704b0dde67ede71a43459ef4486066bc265f0c6

Observation 6fee0551-b613-43b9-967e-82b676493a16 · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.098130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.098130Z digest=sha256:9fa0daa260f2ad9c0579737a56442c8d2a3363311f16ee4705130dbe90abba66

Observation 7b5b40f2-6272-44bf-862d-235f66c09f47 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.103612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.103612Z digest=sha256:490d13cb0ab7ef861a81f38429abbf19407e47e7cae4f5b1506c55a712ade89f

Observation 5d9c1167-71c7-4f42-b00f-5b00f1cab400 · outbound

This paper cites Qwen2.5-VL Technical Report.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.108764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.108764Z digest=sha256:181f74239f977216dbc350f3323c72fa28d04d4063f99806d2fd58ef1a04e31f

Observation 4ac84183-2ef7-48fd-9733-2447b3d9865d · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.113974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.113974Z digest=sha256:e4ffe065dc5f29d451fc6ad403e6cfff83433c7f1a1038cb30f8663f81e93c99

Observation 7007f82a-40d8-4a84-a470-e257ef89e5f7 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.119176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.119176Z digest=sha256:e0f56b54839b88188d697361d8f000e494fa069e69d8a305a6402ca9738324e0

Observation 79510706-cb6d-427c-856e-91a79b51841e · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.123949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.123949Z digest=sha256:395af0ecf51907d59de1d3064094f8fd65776735b4fe953d4896b1c197541a9d

Observation 3498b3d3-2722-46cd-8122-063e95bc77d0 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.129002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.129002Z digest=sha256:8a13dfbdd1cc0018643be4621c0d1657eacbe006044260a757bea7420b7720c4

Observation 2e91de8e-30fc-4eee-b29c-7cbea4208569 · outbound

This paper cites Don’tLookTwice:FasterVideoTransformerswithRun-Length Tokenization.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Don’tLookTwice:FasterVideoTransformerswithRun-Length Tokenization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.134194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.134194Z digest=sha256:4c22e08ed93d08deb018330051d846a1075caddccbac16faba33f8227e7da5b6

Observation 4f5669c0-2d76-4292-8a3a-f65f1538cd52 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.138543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.138543Z digest=sha256:a4a2ff90a15acc4f0e2b359eac0f5d6ab3b019b9cb1bd7fcdd7973bb77c70719

Observation fc077c45-0c18-4a3e-969c-81c362576fd5 · outbound

This paper cites Streaming Video Question-Answering with In-context Video KV-Cache Retrieval.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Streaming Video Question-Answering with In-context Video KV-Cache Retrieval

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.143082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.143082Z digest=sha256:63304c978965003f16c65ddbd2f4655856378c6690c899d33419de26a366eb75

Observation 136859f5-7221-4d5e-b454-337be4e9179d · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.147664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.147664Z digest=sha256:106bad54e68aa7d27f308b7dc8cbec0a6496271c359040c22de021714a2faa69

Observation 271c074c-959e-4b9d-80a0-4d6314b54eef · outbound

This paper cites The Llama 3 Herd of Models.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.152077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.152077Z digest=sha256:a31dcfbad9e41468bba940b59ebd24ed1af3f649f68547320a001ad264b68117

Observation a0dd4c32-5709-4aeb-bb97-694de14e23d8 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.156824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.156824Z digest=sha256:060876e25b03270f1dcc2d5da1fa170e59fb5207c6d4c8b5fc4bb38cb2d05723

Observation 2802569c-fcd2-459d-a41a-4a35c42b31c3 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.167772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.167772Z digest=sha256:8aa643995d0cb2bccbe3980295c617128b41386cdc65a5a334d737f4c9ae3ad2

Observation fc1e19d3-53d3-4cae-9bf8-742d4c5b998b · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.173037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.173037Z digest=sha256:3358c13597a968e2964f1fe26ee3a974d774671962ddf2a5419adc3c971a18b9

Observation 3241c7b6-b5e8-4d60-ac65-25d3204c055e · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.177981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.177981Z digest=sha256:93599fb09ba4d9bffd6971d555ba44a38161e9906df57145a7d32cb585559075

Observation cf4d288d-e789-4cb9-9982-8a91b8b4171d · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.182670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.182670Z digest=sha256:47c5d1d9bc66c59d6b16fa12bd77c467e13e15bf3e44b1013b28a602d565e0c2

Observation 5360a405-d82a-49a2-9271-b8addb30d70c · outbound

This paper cites Multimodal Pretraining for Dense Video Captioning.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Multimodal Pretraining for Dense Video Captioning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.187274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.187274Z digest=sha256:f6a8a25c0a3244ba2a71ecc724fb190ad10697d11c231ab38b5c05f2216a7953

Observation b394037e-7de9-474e-b409-151741d0fbbd · outbound

This paper cites Online Video Understanding: OVBench and VideoChat-Online.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Online Video Understanding: OVBench and VideoChat-Online

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.192617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.192617Z digest=sha256:d9c8f4e8cb2c0e072649ddf3a2171c6cf80af1464ecbccd6caf2f935e6d8dc0e

Observation bf94d745-1f9a-46bc-9f3a-71be881ee060 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.197269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.197269Z digest=sha256:bfc4965eb4d9412d21cb352562e642acf6e5330d0eb5796f0d2839d3ae33a5b3

Observation 64b3f170-e535-4655-88f8-119a944dfc66 · outbound

This paper cites QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.202370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.202370Z digest=sha256:e79736502ab21ac89090c7e357827a205e01554690fe78d7bc7b2b66eb443d76

Observation d19ef454-21c6-4bc9-9061-57d80a06e500 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos LLaVA-OneVision: Easy Visual Task Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.207400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.207400Z digest=sha256:734018afc50cf8c05801df189ff41c5f57525643217c0592dcd48fedf4228e1d

Observation 107b2991-8980-478d-8e55-612140f73ec1 · outbound

This paper cites LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.211905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.211905Z digest=sha256:eae0bf1f256e369a03d5b06d03368e20ebf4b2635ff2ddba407609c81611d619

Observation c4d8983a-6629-4941-bb5d-a0bee40abdd6 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos VideoChat: Chat-Centric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.217162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.217162Z digest=sha256:8d8d85af804653d59ac935bb73890a51f29388152db42cbd9203f3e8771ca886

Observation 26940ab4-974a-4225-92ce-84f720a61896 · outbound

This paper cites In ICLR 2025.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos In ICLR 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.221993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.221993Z digest=sha256:7806d12ab347fc3a9320aa20d58df936dede2e6d9ff47c9e49a1c300c8646d8a

Observation ef6225a9-3859-4964-aab2-ccc56aaf5b35 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.231104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.231104Z digest=sha256:8a8e12166fa8989a8b892f8cb7702c551415a8b800bd3fc77b51328268bd3cac

Observation e0699005-bd32-4f9d-ba5f-281e2af5e08a · outbound

This paper cites OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.235208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.235208Z digest=sha256:ebebfa34193ac6cf85bb6b81a054255e41f3eaae5c284f6f88adfdbadd37671b

Observation 8d0739f2-593e-48a2-979e-b2d4127b843a · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.244381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.244381Z digest=sha256:8a8d5f9022141e56c58945a923a320ce44a91fd2b5f33d5effdf855431d56a30

Observation 63b43313-834b-48d5-8502-d609831f2d14 · outbound

This paper cites StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.248339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.248339Z digest=sha256:0f0d3d768f932518cc92345f00ee225b36d6415ea1225b1b15b00f8dd9824a8f

Observation 971356b9-e82e-4666-aa5a-14d6f6f9974d · outbound

This paper cites Vila:Onpre-trainingforvisuallanguagemodels.In Proceedingsofthe IEEE/CVFconferenceoncomputervisionandpatternrecognition .26689–26699.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Vila:Onpre-trainingforvisuallanguagemodels.In Proceedingsofthe IEEE/CVFconferenceoncomputervisionandpatternrecognition .26689–26699

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.253023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.253023Z digest=sha256:36a053ce5c3cb9fdf0434a0f9dc15b79765b34605fff14166e2f9cd3461ea26b

Observation ae55f94a-be35-4c4c-9487-139c830393e3 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.257196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.257196Z digest=sha256:089b64a2af5d13c7400576a807a7b3dfea7f4341689bfebfde6a33e6571ee35c

Observation 87dae969-c4b4-47cc-b626-ee06ddf7faae · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.261466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.261466Z digest=sha256:bbac90a2e1086458e25eb87f981fd28d464245e1c1cdd78e159715b36a02d5af

Observation b479db0c-347e-4581-8db3-d260e003ecb7 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.266006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.266006Z digest=sha256:8aeee65a5095935f7ccd09ec49ba6f8683098ac42f69079c3991683385987d17

Observation 44e18775-1478-4541-9ebd-ebb3f46a1cb1 · outbound

This paper cites Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.274835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.274835Z digest=sha256:b0f4eb6f1f7bc15798f27c4cc9fbc04183045fb7ac1b573720283c7e42d15296

Observation 5bc27a61-04d8-46e9-8af5-bb45ed83bb6a · outbound

This paper cites Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.279636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.279636Z digest=sha256:d191b99d225e6e93ada1b1e3c257ee0b2209a32babfe8efc68c0b0805f2cc5c6

Observation ee3f9cc1-b145-44a8-af2b-cd83b855f63d · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.284240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.284240Z digest=sha256:61d53e8a02e4db0f4d89243f7a9bd011f36f13245434e648ff2a1e89d97182b0

Observation 43fce36b-8251-45e1-92f6-b3fde6fbfea7 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.288786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.288786Z digest=sha256:2956fc183be5654fc32993268fd0467d6851ce1914a3e18ba3811ea8805706d9

Observation e27bea18-4da4-4852-8060-5857443d36e4 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos DINOv2: Learning Robust Visual Features without Supervision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.293093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.293093Z digest=sha256:9824f2a747320b57ac421de7a12f57dcc43ae6e9847bea848040228192b2ab3e

Observation dacafc79-a663-46f2-bd74-010f335527cb · outbound

This paper cites Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.297623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.297623Z digest=sha256:3b375320e0542bebc10fd76dda589c2eec2ff7d7ae1231ee58e3016f603a7814

Observation 69f75cd8-91ae-477d-8580-404c88efe4ce · outbound

This paper cites Streaminglongvideounderstandingwithlargelanguagemodels.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Streaminglongvideounderstandingwithlargelanguagemodels

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:29.133826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.302174Z digest=sha256:14269e419965c8c5624d9a35b97c2a52261d4fe38dcf154e0e602c383e80d594

Observation 562f1fdf-064f-4014-a903-d8a630884241 · outbound

This paper cites TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.306650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.306650Z digest=sha256:8d9e37f3fa39b4135c6a9f616e89f9d95ca643b4f6733dd434da967dd6857757

Observation 60541043-5b6d-4494-a1be-439780edce04 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.310930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.310930Z digest=sha256:a5ef39d5d36cdd01851890adddeb7d9f5ce2c4fc1837c5fcde2cec19dd2942e0

Observation c5409593-b7de-468c-ac8b-d76472e8e01f · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:29.108014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.315586Z digest=sha256:3e7820baf4ef00f03d8cf0e73b7fa21b5b5bfac3a9aeb466f16ff6ea1546c850

Observation 7cf0d026-9b17-497e-b7a5-daf9a3741aaf · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.320182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.320182Z digest=sha256:d37e07206b6780592a4b0389d1ba1a2e982fdfa7fe09a7d5d3bbaf9910e2b497

Observation b352433f-0ce4-4808-a468-4c735059d1c8 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:29.093257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.325098Z digest=sha256:c7550c0dc821651dfcfc3946aa647e3516115315ad82fe0bf3089f34c545bde3

Observation 26a2d857-9ee4-44ed-821f-7f1c35c99eee · outbound

This paper cites TVSum: Summarizingwebvideosusingtitles.In 2015IEEEConferenceonComputerVision and Pattern Recognition (CVPR).

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos TVSum: Summarizingwebvideosusingtitles.In 2015IEEEConferenceonComputerVision and Pattern Recognition (CVPR)

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.329403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.329403Z digest=sha256:9895c5e79a443e41e093e09352c9eecc15bc0782fe78272ac22854087b92520a

Observation 845c9174-609f-4f35-baab-bdb5e3309ced · outbound

This paper cites Iscosine-similarity of embeddings really about similarity?.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Iscosine-similarity of embeddings really about similarity?

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:29.078801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.333522Z digest=sha256:41b99b36c3fbf80431d47d9f3cad57fe477be2e8b5126b380a20fec4a40f6b73

Observation e2ef9b46-e4c6-411f-898f-bd6695096f11 · outbound

This paper cites COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.337707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.337707Z digest=sha256:a5e663ce08bc4fb7a7d408220fa3fdd6f25a4b825a7307a028bc17e3e5e07127

Observation 93679d75-7309-4972-93c5-ec47c6176dcc · outbound

This paper cites DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.342314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.342314Z digest=sha256:27c808ef08a93801e274cc937575b33acf0ff825a562bda715286099b3604a90

Observation e843b2fd-62ed-42d7-b04f-c7bf6e16da86 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.346774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.346774Z digest=sha256:4f8ca4dacaaffae4941f341036032d030f40e25edbc389d0b18141425dd95f45

Observation 9dd6115e-0687-4509-bae5-2814be6b5d52 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.352149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.352149Z digest=sha256:3be52cef2885c9d04a684bdca7f49de85277c82bac50408dfe1e0aec347b88aa

Observation 363671ca-fd42-49a6-820a-314a63e3b0a7 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:29.063219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.356608Z digest=sha256:29f48857f0329a2de0113593005a729b10c115765529f248c5f9aa91e601f3f5

Observation 2eb589c2-a5ab-46aa-bfd5-4ffea80d4470 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.361147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.361147Z digest=sha256:facf8d8413c1f6724253358d64b33b1bb2441b8460e8d1ef7d4424a520f8328d

Observation a95bde9f-f54c-498b-a23b-4bfb19bc62b3 · outbound

This paper cites Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.365910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.365910Z digest=sha256:ef8b50abbe3412012ab539ca9c224ca4b135ef4658aa1bbe7d8a78b205ac671a

Observation d3bb6f10-c488-4529-aff5-d79133254972 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.370415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.370415Z digest=sha256:7434828cfba2f876766e8ebbf820452e114134805cfabdc2009dcc2dacaceb73

Observation 1e29ee4d-3e99-46e4-96fa-d91a28e12452 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:29.046512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.374807Z digest=sha256:b50a2d05bd52d1731459bfd326cf90cc27d1e773781d6a7f19b621fce0a46235

Observation 50934a88-cc54-4766-b40b-110efda1ead2 · outbound

This paper cites Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.379932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.379932Z digest=sha256:76f0ceac4ee5e17c0ccf51427872a1c9208553e889ddad2d27ad550b3ef72fed

Observation 0b1a58a0-8c95-4b50-97c1-074d47e3ba51 · outbound

This paper cites Qwen2.5-Omni Technical Report.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Qwen2.5-Omni Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.384477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.384477Z digest=sha256:d5535d477dc2bdb0237def0a0734b0d16ddef210b968365c622a4a0602607b2b

Observation 5077567d-e7e0-4f0b-96f3-5b718854d4d5 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.389581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.389581Z digest=sha256:531650f42b6798e4b5a2ec71074b03fed6ba2fea4668a3f0fbc571ca98afee5a

Observation 89ec4016-9523-4b04-9527-f730c20d6efa · outbound

This paper cites Advancing High-Resolution Video-Language Representation with Large-Scale Video Transcriptions.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Advancing High-Resolution Video-Language Representation with Large-Scale Video Transcriptions

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.394273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.394273Z digest=sha256:73974d1d4cafb3e6fbe846f9518cf64134322fb86af40009ee434d3a79f20f44

Observation 6d393e25-0515-4edc-aafb-8cd701e3ff81 · outbound

This paper cites Qwen2 Technical Report.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Qwen2 Technical Report

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.399021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.399021Z digest=sha256:b5db53ab7794a2f39d7916f3835bb658bc4d5e4aa8b25e98b4cae0f17d37f11d

Observation 57884cf9-3f0e-4026-9fcf-f4b9369a722d · outbound

This paper cites PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.403604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.403604Z digest=sha256:9831e7a89752a0107c2d0e1fd7712305efb56f08bfbd74ee51042bd9ac574755

Observation c2d0699b-afe2-43e0-bcf2-3868504c7088 · outbound

This paper cites TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.408169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.408169Z digest=sha256:33aa9bb8dee4e6fb8606774c2887503d5d0732df9308ec78b9dc560128bea524

Observation 2b93f80f-18c6-4dcc-b3da-46e357c340b1 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.413010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.413010Z digest=sha256:0c2cc30f3afc8c6285902a308e5185e888f9e524af3a0bac305ce643a553c36b

Observation 15349772-bd33-4549-8eea-4463566cb1df · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.417308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.417308Z digest=sha256:a41ce4aa1b3ace69d8dadd999a40e431e20e682c4fc8bd1438c02d1a423305fc

Observation 7ce5fbd6-9fdd-424c-9dd6-afb05a8886fb · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.421997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.421997Z digest=sha256:b372667adbcd671e0c2e2f17aff1cb8bb8a2b74040f3f530726e99b5acc1d5be

Observation 2e1821e8-d26a-4325-9b71-2e41af93635f · outbound

This paper cites Movie101: A New Movie Understanding Benchmark.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Movie101: A New Movie Understanding Benchmark

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.427239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.427239Z digest=sha256:5e3096ac7deb9ac7770837f7577e105ff64dba48225d50f338c19e61d3fc3caf

Observation 21c3152c-a3d7-417d-926c-827d6d988ec8 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:29.030799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.431951Z digest=sha256:457f3f1be03db5eccf3df32e87d6a3092387bfa7913db5620a3ba3650b9a6b47

Observation 9c628f04-1a0f-43c6-b7f7-b4af476d59e4 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.436518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.436518Z digest=sha256:cf000081a4090416ddd91cba8a2ec75a57098d170ccb938d31e041b55a962aae

Observation 9071930c-4076-4d47-aca5-325f7cdc98be · outbound

This paper cites Video-LLaMA:AnInstruction-tuned Audio-Visual Language Model for Video Understanding.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Video-LLaMA:AnInstruction-tuned Audio-Visual Language Model for Video Understanding

Reference 74

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T10:47:29.014354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.441111Z digest=sha256:fb0cbd0397d7c37f71b4fed78ce8107c8f8c5566594d35595043a492a82a83b7

Observation 30fc1c39-01d4-4181-a788-ec5b79a5d8b9 · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.445344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.445344Z digest=sha256:00dc1aa5fe56fccf64cd14fff36ba4ea82d88479f9abce9b368696ab49c0b7f4

Observation e1a55fc5-d6f7-4e78-b033-919f27341b42 · outbound

This paper cites Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.449972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.449972Z digest=sha256:4665744e7e7db361f293f7aad79e52dbd293969140b4f9828d0e312475390582

Observation e4f1732d-a2a3-497c-926c-d717722e1696 · outbound

This paper cites InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.458880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.458880Z digest=sha256:a21f8863b1eefac951c4674c1be1a84027810e19833b6a1fbcd046f016339199

Observation 7f5f2872-5095-4f7b-a7df-799cb4c2d40a · outbound

This paper cites Long Context Transfer from Language to Vision.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Long Context Transfer from Language to Vision

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.463436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.463436Z digest=sha256:5d3ce3c53c002729e4e81908db67e6979f996e256114f7f22e16b32ad1ffb164

Observation f6021fd3-3cfe-4b2b-b3f8-d410c9d2a54d · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.468229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.468229Z digest=sha256:a2d62294bc789828a96f0edda3a1ed2592075e7a8b7fa7961f102106e3f59510

Observation aa139f51-bedf-4e4a-a790-677f64c4ab91 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:28.998891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.472582Z digest=sha256:e2bb6fa963e5dcb14799f442c9edf410a4545dfe03d11b6923478c2a250e2e15

Observation d8daf59c-8526-491b-9dbd-0689dabdc2b1 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos MLVU: Benchmarking Multi-task Long Video Understanding

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.476958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.476958Z digest=sha256:f7ecc5c37cd5dd4d107f22b9bd87f2877ccf25212b4425bc7a0f2e0ba1e4024a

Observation 2c637bb6-c0e1-4365-aeb2-b2ce99df42c8 · outbound

This paper cites What specifically did the woman in red do?.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos What specifically did the woman in red do?

Reference 83

Resolution
verified exact
raw_fallback, observed 2026-08-16T10:47:27.742015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.481623Z digest=sha256:2bde7e8857ee90b31c9c30b75935b1a44207583c739a6e318cd8d0161ecc22e2

Observation 5dc0a291-a9f7-4628-9189-4c347f536942 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:28.982946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.487430Z digest=sha256:99d0939df01dd051442d378ac7da86116ccf739c5144e48424022ac61100a69b

Observation d6eaea5e-afe6-429f-8904-6d0a2612da66 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:28.968516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.491754Z digest=sha256:8a8cce0f6bd9bac6f30c50bd42975baa738ade3e36217a8bc28700cb010ebb4d

Observation 1468b03a-e953-474b-b861-c6c856c55aba · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:28.954233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.496136Z digest=sha256:9681e60996551044a70fd15512c86c841ff440557addb2460340a744abb46373

Observation ced68de9-e08b-4081-a72d-bd68d67b5088 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:28.938683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.501279Z digest=sha256:efba7faa62b52e3e638a6fa2115acb6029086e3843f35e03f61c5389200d8550

Observation 7e3bac44-c3d4-465d-8ca2-0103fc95d03e · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:28.924923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.505800Z digest=sha256:4e8ec739a16cdbc5c99153775dbf81080b040dbc6e7a956158dae27fcf184d7a

Observation e48a6b4c-1d73-4729-9ef1-60a2868867e1 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:28.910631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.510130Z digest=sha256:6a23035d8a3434ac448b0adda086f437c0f1e77639e848f79a9bdcb25423f5e4

Observation 3039ca9c-81b2-4b41-a577-32be2eb1fc4b · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:28.896498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.514545Z digest=sha256:9dda042bb5f3d9c13c9f1cee7df2d6675f3ce5416719b6c14c908c0e32886b7b

Observation 2732f0f9-3674-4ef6-8e83-abdb1de8970e · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:28.882122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.518872Z digest=sha256:0cde0ce225344f1fa214d2165e7087262ec6d50f6792b64e2fa704974d3cdb73

Observation 7e3784d1-1663-4669-a1f1-0c66b9cec03f · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:28.867776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.523091Z digest=sha256:87c394e2869d3ee649b16f0da5cc9c2d6322984d72ea855dab8d805cebfdfc9a

Observation 2425f959-a61c-4b00-ad5e-37cfb23e3b2a · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:28.852937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.527429Z digest=sha256:cb0b99c8aa77b26a332099b4d7f2d2eec6f5a7d410354ad132b04f27279791a6

Observation c5b94f0f-d3da-417d-963a-ade005ceec62 · outbound

This paper cites an unresolved cited work.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:28.836442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.531851Z digest=sha256:1ceb8261c0b44e8f81480c6c8b1d5a991ed9423ef00b36d47a1343e07ec5c98f

Observation 79ab746c-798b-42c2-8208-ebcd573c5ca5 · outbound

This paper cites In Frame 201.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos In Frame 201

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:28.821895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.536363Z digest=sha256:abd1711169f7548148671a2eab57da9b87564da3528fd088599f3d4d18dec33d

Observation 02088992-95d2-40f0-9b7c-174722b03431 · outbound

This paper cites When the person was in the kitchen.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos When the person was in the kitchen

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:28.807224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.540878Z digest=sha256:3b014fbffb670722237d0e9914726c3ffafc9a6b6127d63dee2ace1462aabe96

Observation cd0ae2ca-c005-415e-a6fb-d1f94523a7c7 · outbound

This paper cites At 1:20 in the video.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos At 1:20 in the video

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:28.790976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.545250Z digest=sha256:327b24d8b3f5a5c139d8d5f3e0275eed0fba69e14ee7edbe4d87fb86d8a3da68

Observation 38d34bd1-3ea3-43a7-ad4d-ec6e13aab2d3 · outbound

This paper cites question.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos question

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:28.776073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:27.549539Z digest=sha256:39a0c4747538b6ad729c3640d6f3ee61dfbd4da10e75d31025a8bfcb6f2efba5

Observation ecbbe075-d80b-46fb-9e5a-35ef8b9985ee · outbound

This paper cites InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.162648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.162648Z digest=sha256:aa6cfc844860f0a5d18b33118b744e22a33d7c3617a10364f31573dd7eb0aca9

Observation 252a97bc-0ba7-4599-a60a-e63c344bbb2f · outbound

This paper cites Advances in Neural Information Processing Systems37 (2024), 32076–32110.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Advances in Neural Information Processing Systems37 (2024), 32076–32110

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.270496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.270496Z digest=sha256:94584b47adeef1b87b72916696d9f146bd3dd65a88199389658932fdf435e45b

Pith citing papers

Observation 46f8794c-4d6e-4420-8c39-7195e583ed0c · inbound

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models cites this paper.

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.040372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.040372Z digest=sha256:9809fa4daa94f0b9cde16b66c1f981f7d06c7ac22477bf842acc24ada5b13c40

Observation d938463f-64c3-4ad2-8bb0-0f1ac0dd97a6 · inbound

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent cites this paper.

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T15:38:44.871097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:38:44.871097Z digest=sha256:66b296b7b50e490bf1629b4ee1408b453852109c93fb6a51152c203f6530adc4

Observation 61ad33aa-d5bb-43aa-bcbf-ce181d58f477 · inbound

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance? cites this paper.

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance? TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:39:06.196486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T05:36:09.208754Z digest=sha256:de1b13046b7082cc8d8142a53ec4245a79e44ceaf2bbd0308995bc5e86d38be6

Observation 11abff98-bfa7-4491-9d0d-7670db5d6b1e · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:53.845048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:a25b510010c1526261138393f0a5063648efdcfb1500cfa5d0544df29da3f468

Observation c6459779-3754-4260-8941-caf8d2ac18f6 · inbound

Don't Pause! Every prediction matters in a streaming video cites this paper.

Don't Pause! Every prediction matters in a streaming video TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:18.125769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T04:32:01.379605Z digest=sha256:ca8bdcaa5f59e0199c09e517c0cab266bf0b847e02b3aff4d7aef54470253f02

Observation 9a56d4c2-d23a-4179-84b3-a42f9c4c4bc3 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos

Reference 164

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:14.311583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:14.311583Z digest=sha256:84417401e25886e26c7c21cde3732d8401043d1d838ff35e9cf0aaa09841f944

Observation 9f90c176-5427-4d24-8ab0-82d7e6cc4a71 · inbound

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding cites this paper.

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T03:21:48.087769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:21:48.087769Z digest=sha256:2ca47df19d8f4ae97456bad797e5cf7929c61e3aa4d7f0799fb223a802c1c4c7

Observation 5571a5f0-8880-4e9d-8541-74175d42959c · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.620346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.620346Z digest=sha256:38a18f28ac9e9477b0d79119fd465f0cfaeba9748be1ae72872111535ccf246c

Observation 0d082abc-cce4-4720-aa24-69bd9a10092c · inbound

Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation cites this paper.

Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-14T04:39:27.373948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:39:27.373948Z digest=sha256:c52f36938972d4f637ba267b3e7ff7b7fff323136d104717602a59a25b34f622