Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T10:49:52.282672Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 11 inbound Pith citation observations for arXiv:2412.16117.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T10:49:52.282672Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:11:49.166898Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-15T19:16:31.884188Z
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a8ed0265-8631-425c-94c2-ad34909d2cde · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30473d58-f116-4053-bd1a-05e46582f1a3 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Gpt-4 technical report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 24cf1b02-3b0c-4215-83c3-b0644fce1bb7 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Token merging: Your vit but faster
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d9f85c15-d746-46a3-9c7a-7fa9c099bb4a · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9b5385b5-af5a-4a49-a855-7686442ec1f5 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Study on density peaks clustering based on k-nearest neighbors and principal component analysis
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 88ec6223-67ae-4794-8672-c06896802dce · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Slowfast networks for video recognition
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80bb6965-da4c-4a9e-bcc2-b4610da4b881 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d3366fe1-9738-43b6-bced-577444a029bd · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Visual hallucinations of multi-modal large language models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 09420e7a-1b5e-45ac-89f0-207dcf335507 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Chat-univi: Unified visual representation empowers large language models with image and video understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 362e2a01-8708-4b6e-8e27-43cf5acbc5a9 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models An image grid can be worth a video: Zero-shot video question answering using a vlm
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f7a02555-cc91-4288-892d-9518d6300505 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Spvit: Enabling faster vision transformers via latency-aware soft token pruning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7366d366-22a8-42ed-b489-fff6b0b1a64c · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfda57c7-00f7-48a7-b788-59aab23c68da · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Videochat: Chat-centric video understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 53e4dbbc-d041-417c-8a3f-fcd9abeda59b · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Unmasked teacher: Towards training-efficient video foundation models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d4ac5416-df6d-4fd7-a6b0-bcab107d6709 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66e32273-bb97-4973-9e37-847d0845211b · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Llama-vid: An image is worth 2 tokens in large language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 88a193d4-94dc-420d-8cad-ec0a1bf467aa · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Video-llava: Learning united visual representation by alignment before projection
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 01479777-00bf-465c-8921-da0de4ad68d5 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Vila: On pre-training for visual language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3d365085-e043-45d3-81b4-72357ab214bc · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Visual instruction tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c1dfc92f-58b0-4cad-acd9-d52448e12134 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models St-llm: Large language models are effective temporal learners
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 444af41f-3654-4986-b6ad-dc803d627f54 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7fdf3f4f-fb51-4c9a-8f97-5b040ac77ae1 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Efficient inference of vision instruction-following models with elastic cache
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b3513bb3-1bba-4acf-bc75-9b61ca1dfdcf · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d64665fc-a78a-4d1d-a015-e1d779b05a8d · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Egoschema: A diagnostic benchmark for very long-form video language understanding
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 655f4a8a-6c93-4915-94a0-32464f0e6001 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Learning transferable visual models from natural language supervision
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88aee4a6-83c0-421f-bfdc-420889855522 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Dynamicvit: Efficient vision transformers with dynamic token sparsification
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a8e3060c-1a14-422b-a9f8-33fec2064078 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ed3450b3-cc7c-4344-87f6-d367817562d2 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Llama: Open and efficient foundation language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b6eb80ef-3fc5-4d0b-a3d0-31a93a4dcb93 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Fastvit: A fast hybrid vision transformer using structural reparameterization
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5427a9af-9822-40b5-8e07-c113af179731 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Look-m: Look-once optimization in kv cache for efficient multimodal long-context inference
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b16479fe-e55b-4b5a-af37-f9777b044479 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Tarsier: Recipes for training and evaluating large video description models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cf22f3e6-e6de-4dcb-8634-df72158be49a · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Actionclip: A new paradigm for video action recognition
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8a6da8fc-f78e-4745-a4da-4d737627cc13 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Internvideo2: Scaling video foundation models for multimodal video understanding
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c16655bd-aa13-4471-99a6-9f99870feabb · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Freeva: Offline mllm as training-free video assistant
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 31e1dacb-9f1a-46bd-8f4b-ecb579dbbf47 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Pllava: Parameter-free llava extension from images to videos for video dense captioning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b2381bee-dfb4-4b3d-bbe3-d858b3a27169 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Slowfast-llava: A strong training-free baseline for video large language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b5078aeb-6a01-4692-9ec9-08b0ff3b3258 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Qwen2 technical report
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3dbe2622-126b-46c3-adc4-dd4a1693e392 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Video-llama: An instruction-tuned audio-visual language model for video understanding
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6d371b21-56f9-475d-bedf-703558c711fc · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models H2o: Heavy-hitter oracle for efficient generative inference of large language models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac1d57e7-0fd6-4e11-b1b8-66fc4376ca5e · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models @esa (Ref
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77aa930a-754e-4b4d-8eee-8428bbef4323 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8357fddf-4eb3-483c-8862-0e1b702a7478 · outbound
PruneVid: Visual Token Pruning for Efficient Video Large Language Models Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1096a89-6bc4-4e70-b5cf-9ab2e5fed281 · inbound
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d544c777-c601-45fb-a23e-d49639e58b38 · inbound
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f29c4805-84c0-4a34-92c7-6508dc49d3ec · inbound
MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 747571fc-405e-4119-b757-877208d82e2e · inbound
TrajTok: Learning Trajectory Tokens enables better Video Understanding PruneVid: Visual Token Pruning for Efficient Video Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 59c2e359-e839-46af-bc6a-e64c09bff64c · inbound
TrajTok: Learning Trajectory Tokens enables better Video Understanding PruneVid: Visual Token Pruning for Efficient Video Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 352f6555-5169-49e5-ab23-c28f34d52a3b · inbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models PruneVid: Visual Token Pruning for Efficient Video Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bf580084-5077-475c-b446-9c4c0ed3c70b · inbound
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 391ab039-e668-49f0-92e4-cdc6a8d5e4a1 · inbound
Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation PruneVid: Visual Token Pruning for Efficient Video Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6037976b-5fc2-4d7c-a168-9c15f12ab2bd · inbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models PruneVid: Visual Token Pruning for Efficient Video Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c93e8ec9-6253-41fa-a4c4-35e752a53ea7 · inbound
PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9ce454e-6fdf-4a22-84c6-d681f1f16dce · inbound
PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.