Pith. sign in

Paper Citation Record · LEDGER

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models

As of 10 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2608.01644.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01644 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:51.667788Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f0cac426-2ff1-4169-b70c-0e2061e4f20d · outbound

This paper cites Token Merging: Your ViT But Faster.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Token Merging: Your ViT But Faster

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.531752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.531752Z digest=sha256:8a019bb01c1c35185c47b13123ebc756239a2674ced04f2f90afa942376988f2

Observation e95ad9f4-fe05-433b-9d6a-1e36fc9ffd25 · outbound

This paper cites DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.535123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.535123Z digest=sha256:bd204e44824445ae7c2a5edfc3eebe344b5f396def2272dd35d5a61a40c107c5

Observation a5070a4e-c3c1-4be8-9a28-ba259dfc7587 · outbound

This paper cites Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.538242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.538242Z digest=sha256:fbac69c2f54309f64518588f360be1b1d9dbb672023e27e4657c110281ad5113

Observation 7938faef-9d75-4293-aa42-e02aab49098f · outbound

This paper cites Johnson and Joram Lindenstrauss , journal =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Johnson and Joram Lindenstrauss , journal =

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:52.644192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-04T23:46:51.541420Z digest=sha256:a4e63cf1274f5d2453063494f8de0f441297ed523570d15e58141d3cfaa7b576

Observation 6faa0741-c26d-4918-945f-2a3936bdff27 · outbound

This paper cites Database-friendly Random Projections:.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Database-friendly Random Projections:

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:52.636330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-04T23:46:51.544042Z digest=sha256:ebf708859419485cbc88548fc979ac9bf8b17a9ce286b64f5d567fc31d237d63

Observation 818cb237-c809-4edf-a4ac-13d5b653a7be · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.546718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.546718Z digest=sha256:d1fbbf1733b4010392088dfdd1a250f55b5703615f5f9f182ad378f1130bb476

Observation 2dad76de-74ca-4d25-993a-6a3554dad9ed · outbound

This paper cites DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.549583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.549583Z digest=sha256:30848ec29c199ce918c1efc16d44536c25baa7aaccb4efb5820d27f19ab4599a

Observation a915563f-32d0-46ba-badb-4d8b2818cde9 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Advances in Neural Information Processing Systems (NeurIPS) , year =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.552299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.552299Z digest=sha256:fea112362e80ebf316ab1f5d7d3351ac2ac4d43565701ed230ce697ea17208c6

Observation 608e3e4a-790e-4780-a5b2-6aeec30f49b0 · outbound

This paper cites IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.554772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.554772Z digest=sha256:ea6120ca3cb37d0106b2979455ea0e4b831974a9d40bc29690e31096a51453e3

Observation 6037976b-5fc2-4d7c-a168-9c15f12ab2bd · outbound

This paper cites PruneVid: Visual Token Pruning for Efficient Video Large Language Models.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.557165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.557165Z digest=sha256:908a565345c8a789166bb95d0da3e40665940e9cd7ac3531d601be8853c92124

Observation b88481f7-f192-46f0-b1f2-d96b29e93664 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.559912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.559912Z digest=sha256:79695faf9ff441baf044c8525560b09ca0cc9777af662fc6e7d493c37e2153f0

Observation 4829ecd5-8452-4bbb-8fbc-339d242d13cd · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.562582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.562582Z digest=sha256:e2edffb5590ccb52bd0629f58a7b49a6f918946f632ba3d80e89cb752a23d595

Observation c4c00660-b607-42b9-a766-4b247d742a41 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.565175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.565175Z digest=sha256:b8222d0f838dd68353d8bad357475ac4fbf33a83b5d1a9ec3331247f585de6f8

Observation 029a34b8-1856-4299-8a32-a5378955ea5a · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.568021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.568021Z digest=sha256:4ecfac35d189c956dde3bcf4d8bcbab3db0122844662a63b36dc8d30fd9cbfc0

Observation e0bfad12-8a05-46ca-a1dc-50349f089c34 · outbound

This paper cites PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T23:46:52.450235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-04T23:46:51.570775Z digest=sha256:c80fccfe5353736744b28fd02520d59197f5afd2fadc8352a8d3808745078ea1

Observation a3c4e44a-af0c-4238-9ed0-9d3ae8b95f91 · outbound

This paper cites LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.573559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.573559Z digest=sha256:820e43f825109be3c2bb077f93642ba9dfdf7f8a230ba7d7a52738bcad7d6730

Observation 9ba06075-77b6-45c8-97a1-15ca8e42cf37 · outbound

This paper cites InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.576340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.576340Z digest=sha256:2202e5c7b5a8d7ad0dda427c6834a3ac1c4eea1e58fcdb7f89a821e4321ce3cc

Observation 123d02dd-81ae-4c86-8b8b-172b6f77087c · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models International Conference on Learning Representations (ICLR) , year =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.580850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.580850Z digest=sha256:7f8fba2cd11b1beb84ff2be3ecc9f9c4cd495478b39fb0368f8412c73bf36d85

Observation cb192944-990e-4e1c-b078-8a1b37e8223e · outbound

This paper cites Transactions on Machine Learning Research (TMLR) , year =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Transactions on Machine Learning Research (TMLR) , year =

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.583327Z digest=sha256:f360280b709b7e1409715701a40f368de01798dbc60a7092329bab4bad74d3e3

Observation 666bc42d-4ead-47fb-92f7-0e53f0c16fcc · outbound

This paper cites arXiv preprint arXiv:2505.18227 , year =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models arXiv preprint arXiv:2505.18227 , year =

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.585864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.585864Z digest=sha256:d4378ed937d0e5b4a4c3f0bb43a22297a0e29e0eb534513896e3405e9a9824c0

Observation 6e0885d0-099b-498c-9184-bfdcca6e8f75 · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.588231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.588231Z digest=sha256:d20baf73067be651ccfbcf6ef9671a20f4bbd6301cd06fea25c2470b2d51d8da

Observation ff6fa96a-a337-43e2-b5ac-823cdd111a8b · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.590927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.590927Z digest=sha256:e2c18641ade9b244fa1618d09ea8fa2e64080227190f70240238f3ab6cac6c35

Observation 5826cc77-e5f4-404d-a2a2-9e5cf40a94de · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models TempCompass: Do Video LLMs Really Understand Videos?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.593878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.593878Z digest=sha256:5cad55c093ee72f4dcf3188b54f4716f1df14a6ab059471826e02b222f8db7d1

Observation 95377740-8756-4f75-bff3-e15c3b62e03a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.596591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.596591Z digest=sha256:f33ed58b60d20300f909493d0a7c4275a2c3a84a2537b23817962509834f2481

Observation 1c803385-804f-4d59-9e6e-d27819625dd2 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.599549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.599549Z digest=sha256:da5d7570929d41dddb32504e37e28217cee5e0ef0d8dc49d9066d48a768d7691

Observation 555bb139-5a0c-4ef6-8d12-e2ea525bb244 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.602320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.602320Z digest=sha256:e4f4ae19dddfbb57d2db4593fedc0bf1240005fc5689b65bae8fa3b6b292559f

Observation 128b6588-ea48-4bbf-b3f9-f3f9cc72aba4 · outbound

This paper cites Qwen2.5-VL Technical Report.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Qwen2.5-VL Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.605205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.605205Z digest=sha256:1e3821aa26d699ee4d50daf5b396cffecf89e4462d0c7d19b6219b9da85ed4f0

Observation 7dc63e75-60e7-488f-827d-82080ef78b28 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.608318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.608318Z digest=sha256:8a39251d6089d75d7cae0d7e5ce3599a254e3c63e0860e6ad8ddb02c36879964

Observation cec09a60-c794-45e3-8958-0ab05708261a · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.611444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.611444Z digest=sha256:cdc903692381acb4cff214fdeecf0dc01ca303896867c0913bc13ed68d045cef

Observation 5c8d3dba-cc1a-4f75-89fc-61f7f0faaa68 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models MLVU: Benchmarking Multi-task Long Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.614564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.614564Z digest=sha256:14d5c31423c72575904f663633abebcebe2ae6d6d99cf6a668efe2d46fe9b03e

Observation bebe9447-190d-4975-8d70-6088898bdcbf · outbound

This paper cites arXiv preprint arXiv:2510.16598 , year =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models arXiv preprint arXiv:2510.16598 , year =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.617550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.617550Z digest=sha256:34fd1aa0fc9bb26134d0383f0bd4e09fe0fde60bb8f157a12b418c54f282bac2

Observation 5571a5f0-8880-4e9d-8541-74175d42959c · outbound

This paper cites TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.620346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.620346Z digest=sha256:7551495bd9a8c858e494af332d75eba00e53256640d47cf99393d0a195a2ceeb

Observation 28dfb864-4a37-4541-a6c0-9505ef6324ae · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.623682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.623682Z digest=sha256:ee0509916066c00bd3df4b1d9ea6c1d5668fae5307d0059030e241c904bd26e8

Observation 444f74dd-a255-4114-908e-61f7da1616fc · outbound

This paper cites Scalable Diffusion Models with Transformers.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Scalable Diffusion Models with Transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.626661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.626661Z digest=sha256:cb6dcc62378259ded9cfa549a4fd0c80490800098703b59a58bb83105737a1c0

Observation ec101824-b7d8-4757-8953-90f3e2cc7deb · outbound

This paper cites Masked Autoencoders Are Scalable Vision Learners.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Masked Autoencoders Are Scalable Vision Learners

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.629660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.629660Z digest=sha256:73e6f51ebc93ffd0fdeae418d502d65a1ca5214dd6654597e57a6cd0417e3dc6

Observation beb3c69b-1797-4c66-853a-d40f3b182606 · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.632640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.632640Z digest=sha256:6515d5cb74aa642814ac26e595e0d887ece0556455c1b2d73ba778745a7206b5

Observation a7fe473c-2c51-41f6-9fe4-6c91d74d1e75 · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.635563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.635563Z digest=sha256:a995d964f4f86b6a97c43bbe8fbe90f9afd85c312dba5c3cbee0813a0cedb5b1

Observation b87acf7c-1da2-4d2b-9cd4-3f9b021cd49e · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.638965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.638965Z digest=sha256:10a5762b03a2a2e4a10d3ab6b968c99a765d5579815e8a32e997d513ecd9e16f

Observation 84d2da63-359b-4e4c-9943-42b6095c3fb3 · outbound

This paper cites DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.641946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.641946Z digest=sha256:bf2f113dec47d5ffb3a80e3d2991b58a4b133b370abb0f3ab9834ea147214e45

Observation 9db99812-1f46-48fa-884e-5eb4287f9270 · outbound

This paper cites Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.645057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.645057Z digest=sha256:c4646bc20b45e8c8368f719b81b7e90d42b585c84d036d9c66744e1a810f5e48

Observation 60ce3a48-5358-4e4a-b3e9-7352a90a19b5 · outbound

This paper cites DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.648153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.648153Z digest=sha256:604b1cfc5cf30eadf781959bc785b4e110f4e2b262c18d18d57a67385ec6d094

Observation a20520da-add8-468a-a26e-38c53cced558 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Advances in Neural Information Processing Systems (NeurIPS) , year =

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.651354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.651354Z digest=sha256:0304a3f4a20473c1654ff3981e7c9c88ece6a85b330bd662afa4e9feb1ae2c1b

Observation d890eb34-3b8c-4d7c-8973-4ab572bb59d4 · outbound

This paper cites Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.654336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.654336Z digest=sha256:38c64678496a80dfa01a692d8737639944dac5614373ad4a862a39a1b9457bd1

Observation 1029c04a-e870-4c34-b21f-3c5564ab4b5e · outbound

This paper cites LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.657161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.657161Z digest=sha256:16920aaa52fedc5eb8cf559c0e54c3796189204b79cdbd5ba30c43a90ceb94b1

Observation b4359393-9c7c-47fa-85a1-44bdbd4b6697 · outbound

This paper cites Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.659904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.659904Z digest=sha256:b7c472fec47cdd6dff2a4e62cef6e0c39a32b1d71b27bd2b12513c0607f70054

Observation 11af31fa-86ea-4456-9369-545eae1505ae · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Advances in Neural Information Processing Systems (NeurIPS) , year =

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.662622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.662622Z digest=sha256:b30207d48a27bd0534a272bcdf852c0d427a9e22d6d8de415f77ab6ef5610cb2

Observation 8cbc6767-9f7f-43ff-b579-578a66df21aa · outbound

This paper cites 2024 , eprint=.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models 2024 , eprint=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.665137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.665137Z digest=sha256:5d1585d491351510f032038b0367ae902db6a113539edd7816a063650efcf9e4

Observation 8b6aee8f-232a-474b-8d9f-7d5811d96d24 · outbound

This paper cites 2026 , eprint=.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models 2026 , eprint=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.667788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.667788Z digest=sha256:7c9167ba202fd6eee1d77d8044ad5195936b08224cdf4b388aed8fbe2ba6d044

Pith citing papers

No inbound Pith citation observations are available.