Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:51.667788Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2608.01644.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:51.667788Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f0cac426-2ff1-4169-b70c-0e2061e4f20d · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Token Merging: Your ViT But Faster
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e95ad9f4-fe05-433b-9d6a-1e36fc9ffd25 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5070a4e-c3c1-4be8-9a28-ba259dfc7587 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7938faef-9d75-4293-aa42-e02aab49098f · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Johnson and Joram Lindenstrauss , journal =
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6faa0741-c26d-4918-945f-2a3936bdff27 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Database-friendly Random Projections:
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 818cb237-c809-4edf-a4ac-13d5b653a7be · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dad76de-74ca-4d25-993a-6a3554dad9ed · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a915563f-32d0-46ba-badb-4d8b2818cde9 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Advances in Neural Information Processing Systems (NeurIPS) , year =
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 608e3e4a-790e-4780-a5b2-6aeec30f49b0 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6037976b-5fc2-4d7c-a168-9c15f12ab2bd · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models PruneVid: Visual Token Pruning for Efficient Video Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b88481f7-f192-46f0-b1f2-d96b29e93664 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4829ecd5-8452-4bbb-8fbc-339d242d13cd · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4c00660-b607-42b9-a766-4b247d742a41 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Flamingo: a Visual Language Model for Few-Shot Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 029a34b8-1856-4299-8a32-a5378955ea5a · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0bfad12-8a05-46ca-a1dc-50349f089c34 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a3c4e44a-af0c-4238-9ed0-9d3ae8b95f91 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ba06075-77b6-45c8-97a1-15ca8e42cf37 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 123d02dd-81ae-4c86-8b8b-172b6f77087c · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models International Conference on Learning Representations (ICLR) , year =
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb192944-990e-4e1c-b078-8a1b37e8223e · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Transactions on Machine Learning Research (TMLR) , year =
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 666bc42d-4ead-47fb-92f7-0e53f0c16fcc · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models arXiv preprint arXiv:2505.18227 , year =
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e0885d0-099b-498c-9184-bfdcca6e8f75 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff6fa96a-a337-43e2-b5ac-823cdd111a8b · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5826cc77-e5f4-404d-a2a2-9e5cf40a94de · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models TempCompass: Do Video LLMs Really Understand Videos?
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95377740-8756-4f75-bff3-e15c3b62e03a · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c803385-804f-4d59-9e6e-d27819625dd2 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 555bb139-5a0c-4ef6-8d12-e2ea525bb244 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 128b6588-ea48-4bbf-b3f9-f3f9cc72aba4 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Qwen2.5-VL Technical Report
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dc63e75-60e7-488f-827d-82080ef78b28 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cec09a60-c794-45e3-8958-0ab05708261a · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c8d3dba-cc1a-4f75-89fc-61f7f0faaa68 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models MLVU: Benchmarking Multi-task Long Video Understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bebe9447-190d-4975-8d70-6088898bdcbf · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models arXiv preprint arXiv:2510.16598 , year =
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5571a5f0-8880-4e9d-8541-74175d42959c · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28dfb864-4a37-4541-a6c0-9505ef6324ae · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LoRA: Low-Rank Adaptation of Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 444f74dd-a255-4114-908e-61f7da1616fc · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Scalable Diffusion Models with Transformers
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec101824-b7d8-4757-8953-90f3e2cc7deb · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Masked Autoencoders Are Scalable Vision Learners
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beb3c69b-1797-4c66-853a-d40f3b182606 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7fe473c-2c51-41f6-9fe4-6c91d74d1e75 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b87acf7c-1da2-4d2b-9cd4-3f9b021cd49e · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84d2da63-359b-4e4c-9943-42b6095c3fb3 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9db99812-1f46-48fa-884e-5eb4287f9270 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60ce3a48-5358-4e4a-b3e9-7352a90a19b5 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a20520da-add8-468a-a26e-38c53cced558 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Advances in Neural Information Processing Systems (NeurIPS) , year =
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d890eb34-3b8c-4d7c-8973-4ab572bb59d4 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1029c04a-e870-4c34-b21f-3c5564ab4b5e · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4359393-9c7c-47fa-85a1-44bdbd4b6697 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11af31fa-86ea-4456-9369-545eae1505ae · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Advances in Neural Information Processing Systems (NeurIPS) , year =
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cbc6767-9f7f-43ff-b579-578a66df21aa · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models 2024 , eprint=
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b6aee8f-232a-474b-8d9f-7d5811d96d24 · outbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models 2026 , eprint=
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.