Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:56.752601Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2608.03083.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:56.752601Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e1085018-dfde-44b0-85f7-78b5c614ad57 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff5ee652-4d49-40f9-9b74-17d6e282f1d8 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c807c83-78d9-46cc-8132-fc2d4d69d817 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2752700b-4ebb-45cb-9490-5e443217e22a · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47de14fd-faf0-4cb9-8de4-a139f1b4bfb3 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Qwen3-VL Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33ee5007-8e23-48a8-99d7-2d9d0a657a08 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Qwen2.5-VL Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 976a2390-5358-40c2-8220-69286d6b0dbc · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Token Merging: Your ViT But Faster
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22c1062b-1245-4720-88ab-07791839a334 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models VideoLLM: Modeling Video Sequence with Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb9c2de3-7353-4897-8dc9-929353f72484 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acc0232e-b8b9-46d0-87e4-a5a1681fad42 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b8d239a-2792-41d6-939e-32466093dcbf · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 036ab384-5f20-4baa-bcdc-384ca0e7152d · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a25a4001-a7cf-4b3b-8750-1cb6cfb6d7db · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8402874c-e6eb-499e-98d4-d89bb27acbb8 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 213f6d5c-f975-441d-bd33-15a5cde08464 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b02549a-428a-4f5d-afbc-31f50c60bb1f · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 165bbe61-23e5-4761-91bc-dfea5340b00e · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0003ab70-5e6e-44b2-a1a1-5853ce217d7b · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1153f228-79c5-4b5a-aff2-6b76bb7f6cf1 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 86489d16-dca3-438b-89c5-24aa30bcae3a · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 70854148-96bf-40bf-93b3-4d22dce6448a · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f182365d-6de8-44a6-adb0-f3e5b6b9c24c · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a631fc30-95a3-4236-afde-8d3068e58e9d · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3e910f7-1986-4a4c-85af-70f08a47fa12 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c9da7ad-f05a-4c11-8056-47f981526d11 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0215bf0b-ad13-4201-bc61-b19e7b86591d · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 775cee6d-21f0-400f-b0e0-ce21db608cea · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ef048d8-0f0a-4a72-973e-fa2a13665635 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c251361c-8b88-451e-85a8-1924fa62f407 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c2ff6b0-2fca-4ce6-9d22-c5e484402030 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b391219-50fa-46eb-9ead-85bf81236d22 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4060da20-332d-4247-9ca4-fd48cd18631e · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9ab596e9-d35f-4d94-a3f3-cf32b1e0939d · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5453f54a-217e-4f0b-bae0-b276527533bf · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 089bbc5a-ef76-484f-9d0b-149371639113 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cfc4506-8302-406f-999f-a53d56ef2b64 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f92f55b0-5846-4249-99ec-12f84baaf23b · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddfc45c4-336a-41da-8c5c-8bd60da3a61f · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c44ac986-dd29-4347-af25-a97e532d0fcb · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7088f210-d73a-4f6d-ac74-d4d8b01b5243 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2073a944-5328-40b6-9520-14fe4d6f4c8b · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models AdaTP: Attention-Debiased Token Pruning for Video Large Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 772fb8d1-53d1-4470-91ea-b2d273c5ff39 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11f3c715-5a64-47f1-bb66-7ac013f47b7a · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a3258b68-4238-4411-aa41-dd5f7abea387 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 64a4e827-a9c3-4bcd-a4c8-082bf7b7ab40 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 694e1842-7425-4edf-8f27-01d056af4ff8 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e2e6ee1-1712-47ba-b12f-9545d5dba291 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 010c3b1e-cb27-49ca-8f15-57d01737d440 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20e58152-9c2d-43ad-8e45-8501ee920757 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 550c3538-91ba-4b70-8cf3-2ab5bec83702 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 75f927b1-301b-44c4-ad50-57b5fb545933 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f8f6d80a-39b2-43df-bc8d-8d8d45ec9044 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11101bf1-83a1-4c81-928b-1444ec23b155 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 001fd366-5c11-4c09-9234-3350672fa44c · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 14e7701e-2074-4b82-836f-dad26b1bd571 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3acfe0d8-6ee5-4625-b64e-e8a1c28917ea · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1759393-e52c-498c-be32-fd0e4b07dec9 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe7b203f-cde9-4956-baeb-2cd35b553471 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6d206d4-edab-46fa-ada6-a614a83cef4d · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e8d6223-e916-4026-a7cf-c291ee1541fb · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Advances in neural information processing systems34 (2021), 13937–13949
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85cc6769-bc5a-4d74-9385-a720e870132c · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2aee47a-d6c7-43da-bbc9-29666f38fb20 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bab3dea0-b8cc-4aa0-b581-ee049fd4b787 · outbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models InProceedings of the Computer Vision and Pattern Recognition Conference
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.