Pith. sign in

Paper Citation Record · LEDGER

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

As of 15 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 10 inbound Pith citation observations for arXiv:2411.15024.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15024 v3

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:40:23.771906Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:08:09.432264Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T14:31:40.739574Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 530f79c1-4e59-4f8a-8606-9bd9662b71ca · outbound

This paper cites Token Merging: Your ViT But Faster.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Token Merging: Your ViT But Faster

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.606124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.606124Z digest=sha256:df97812bdf8bc7f7933d9e397e669e6212d1233b9e1d973bb61c9425a3dcccf7

Observation 7cd564c9-5191-458a-aedd-52c0215a2c01 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.613617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.613617Z digest=sha256:662254123de9d8a2ed84b1d5eeb8a950f6d5460d2a5a42b4de832447f40acc57

Observation a2a99899-602b-4928-917b-a694295f98bc · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.616706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.616706Z digest=sha256:8ea4671f0d5e8f6da9d7a31403e1a80c23998254a5768c176ed7a8026fac375d

Observation df522aa3-0ead-41ab-be2a-79ddb14ba327 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.620060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.620060Z digest=sha256:c90765d4ab05df9e825efd3fbc22c95192a8b55f10b97f01c026e323f0727ad4

Observation 95365fbd-939e-4e5a-a890-886f386ae174 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.623746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.623746Z digest=sha256:7b393f12df27f36efa7e178bfdd0e4ccbdf4260f5491ebd74750c376e38000ac

Observation 662dffc3-0195-443e-ad69-8d197f7c7c51 · outbound

This paper cites an unresolved cited work.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:40:24.201372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.627431Z digest=sha256:d71eaf60876ba851ce9fd212e3ff1b33c712059075153cb99944dc952041f60c

Observation e3850772-27f2-49c4-a1e0-a3a152f19c32 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.629903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.629903Z digest=sha256:ae91a6dc7f1c66d3c5138364e65b8356ab56ffb75698b048ee592535b129afbf

Observation e8b7dab3-7667-4bb2-91bf-6d44d6c44b0b · outbound

This paper cites LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.633017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.633017Z digest=sha256:9177b151afaed15da8a9c339075fd14b7966ca0684491ba3fde1acc0ad36e558

Observation b194b244-12ef-4680-adde-3143b57ddb2f · outbound

This paper cites Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% perfor- mance.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% perfor- mance

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.194332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.636820Z digest=sha256:b50d46dd9f8c9e21cb696a909ab0aeb46454faef9b1a28370e32a1734977a5a0

Observation bbf33c42-8533-40eb-b11c-1c4948d21924 · outbound

This paper cites Masked autoencoders are scalable vision learners.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Masked autoencoders are scalable vision learners

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.186675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.639827Z digest=sha256:1cc399c48f20ce9559ddcdf1ea4d433defd4b96460f3e377c9de1aa60dd78513

Observation 56211308-d450-46b8-92cd-3d234d3530b8 · outbound

This paper cites Lita: Language instructed temporal-localization assistant.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Lita: Language instructed temporal-localization assistant

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.178893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.642664Z digest=sha256:a137885ef295550353f6a88131c12f45c9ab8f5403f3fe6e221b2be8968f9d43

Observation 6bf8a6d5-43ad-4ec1-b6b1-ed857fabb90f · outbound

This paper cites An empirical study of llama3 quan- tization: From llms to mllms.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models An empirical study of llama3 quan- tization: From llms to mllms

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.170570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.645458Z digest=sha256:8976ef6078205b2e02ba8971151e2d391ad926b7a013e5d56ad91391ef0bc319

Observation 0d8fc5da-66ec-44b0-9c2a-62c868714354 · outbound

This paper cites Phi-2: The surprising power of small language models.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Phi-2: The surprising power of small language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.162483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.648110Z digest=sha256:c37d223e2afe72cb16552fff36afaf7f333930196279e56aac7ea572fa672689

Observation 8c8faeba-6d76-4ed4-aa97-f43027f3e992 · outbound

This paper cites Effec- tiveness assessment of recent large vision-language models.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Effec- tiveness assessment of recent large vision-language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.155676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.650708Z digest=sha256:8f05eae80a2d7fe2596678559068cd5ef9d41b8c615a0b5334ddd6ff9dbadd66

Observation 890ac9fe-8c70-4b02-8a4b-12d0d1fe1eac · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video under- standing.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Chat-univi: Unified visual representation em- powers large language models with image and video under- standing

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.148286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.652977Z digest=sha256:8831d424861b094f3908cf0fa59e9fd8feee931421f8dc8ab2b5f4d833b3a928

Observation 80b32cc2-0d2b-4c0f-8c76-96ca4d5939a0 · outbound

This paper cites Lmms-eval: Accelerating the development of large multimodal models, 2024.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Lmms-eval: Accelerating the development of large multimodal models, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.656184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.656184Z digest=sha256:2ee1394f643104d7ed605618eb80a69466be313dc7ba7bcbcae7a775aafc2913

Observation 444d4a0e-ab21-4579-958c-ccb442e8d808 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.659409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.659409Z digest=sha256:582ca8f4b1f066b1b324a547889007ba638e293d34a3be88b6eaf62fabc0dc94

Observation aeb20722-42e1-4cb7-b13f-c73494ca640e · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.662070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.662070Z digest=sha256:241ff47024c79c11fc129c91f103e96db6f654e1ba71945de0aefd5d70f9823f

Observation 4ca789b6-071b-4145-ba0d-ae457bde8cca · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.137189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.665109Z digest=sha256:bc0d3f2938e88e44bd63585cae936ec1535015d742578f1eaa25b832d1e16650

Observation cca5709e-d411-4a1c-baff-a9270c33971d · outbound

This paper cites Tp2o: Creative text pair-to-object generation using balance swap-sampling.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Tp2o: Creative text pair-to-object generation using balance swap-sampling

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.128620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.667682Z digest=sha256:ab4be9136c025b990cfeda34c442116055e26dddfec2cb0a1e475f5dbd42cca7

Observation ad9317d3-9ab5-4df5-877d-82cced281e69 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models VideoChat: Chat-Centric Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.670848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.670848Z digest=sha256:530b3c773817ab16d6581e5f176b6da2b5b3de13d1c43825d527e9e38f82343d

Observation 3ac6aed7-a5d3-4faa-b167-7b296074fdfe · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.121356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.674434Z digest=sha256:61ddf823a6553757c7a09faf5b3285089cd26843df688671a8cae49d33ba03d1

Observation 98ff61bf-30bd-421d-8bda-b5463da11204 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Llama-vid: An image is worth 2 tokens in large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.114221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.677559Z digest=sha256:f95cd908a2efe3fb03bfd58631a713f799b49a8fc707993c75ff0673384a3af4

Observation 686683fb-6050-47cf-a523-1f7a04717388 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.680826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.680826Z digest=sha256:541abc1b9e35fbf641761839a6ff1caa3e27b1bda01432edc7bca85aa36cc870

Observation 40330da5-7f79-4122-a3dc-ed1e6da79230 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.684059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.684059Z digest=sha256:d809b0a1fd2423b14124b6a59ce9b1c4128a850192882899144c308e2902f0c5

Observation 1d988134-9f0d-49c4-acc3-8016818a5067 · outbound

This paper cites Vila: On pre-training for visual language models.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Vila: On pre-training for visual language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.105739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.687112Z digest=sha256:57526fcd4b0387bb0228221e1fcde34e8f98cc1ae74a159fd5cbc15d2e3ac3b4

Observation a101bd37-0ac1-417c-baa6-8b59d8bce438 · outbound

This paper cites Video-Teller: Enhancing Cross-Modal Generation with Fusion and Decoupling.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Video-Teller: Enhancing Cross-Modal Generation with Fusion and Decoupling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.689402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.689402Z digest=sha256:fa307c9df44fedd400200dfac0d7f2ace42e43f8948e04cab33b01fdf7ceb209

Observation 80057844-2fee-4228-a585-4e8065661293 · outbound

This paper cites Improved baselines with visual instruction tuning.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Improved baselines with visual instruction tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.692236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.692236Z digest=sha256:bf2630cd01884b171529e8a0dce5aa18a16bba5741d4424b09e3ff85c7d9d897

Observation c5f046d1-4fe9-479b-858f-621ffe97f0e4 · outbound

This paper cites Visual instruction tuning.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Visual instruction tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.696074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.696074Z digest=sha256:fd1738c34661a53f18ce52713066cb3a1297f6b2e75918b2a9329428d733ffb4

Observation 0e7c4b62-9d33-4e80-aea7-1a1b8a3af018 · outbound

This paper cites Video detail caption, 2024.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Video detail caption, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.092122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.698977Z digest=sha256:ee38baa63686598b961de71a9ec2107497c0fa007e5c8970f671f699fa79c048

Observation c073fbc5-43d2-4f94-954e-6da67bfaee97 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.701759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.701759Z digest=sha256:b0f998d93d3963c385789f271679be6d7893ccf11f15d7aad239a07a124ae834

Observation b75e6fbb-1254-4b8a-8ca3-76666289f524 · outbound

This paper cites Gpt-4 model, 2023.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Gpt-4 model, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.085000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.704662Z digest=sha256:2710e45f4fce9e1ca879e250cbd1dab7217c848b4a0959ca65359aca5fa9e98f

Observation 0aba2f3e-02b9-43dc-856b-01242661eed0 · outbound

This paper cites Per- ception test: A diagnostic benchmark for multimodal video models.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Per- ception test: A diagnostic benchmark for multimodal video models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.706916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.706916Z digest=sha256:6794ce7075117312eb6274d80819c1c0aacc17cc3fb00e990a8692d546e1432d

Observation f0499b5c-2f5d-412a-875a-3c0fa62174d0 · outbound

This paper cites Learning transferable visual models from natural language supervision.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Learning transferable visual models from natural language supervision

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.073626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.710276Z digest=sha256:b33742b79aa0a18376c309428c38e541af77b5bf4622aaf6a4d46034e1646fc1

Observation 03bae0ac-1fbe-4342-b3e4-9420f2fece5d · outbound

This paper cites TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.713555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.713555Z digest=sha256:dbdc24367af897aa06ebba292f9c4bb5189850060bdbfd81a1a6c113e1fcaa0a

Observation cd0dbfeb-3365-4a3a-a75e-4a19c42f8261 · outbound

This paper cites xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.716671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.716671Z digest=sha256:14280707bca179e8b3143b9f191d2d02f319112447610bdd19538bfdd00095f1

Observation 7118a95f-454d-4822-be9c-bf4500cacae8 · outbound

This paper cites PB-LLM: Partially Binarized Large Language Models.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models PB-LLM: Partially Binarized Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.719444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.719444Z digest=sha256:6f8158aac26e17268e457951d660ca9e19dcfcd43739e8c5b83e985bb47a0b4c

Observation aea40c27-f9b4-47bc-8ff4-03947e4ffb37 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.721994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.721994Z digest=sha256:2f7f7c622363cd9e13abb476dfacce41966f058c9c69405267251f0fc0cb9c90

Observation 3b10fefc-2f69-4e0c-ab17-c914a975ecb5 · outbound

This paper cites TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.724449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.724449Z digest=sha256:7a3e78574f6a4e79b89467c678de0f751990005aec9109ab9d4d2a398832dc4b

Observation 2a70519a-1c22-4609-9647-d7ff754b9db5 · outbound

This paper cites Video- mae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Video- mae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.065685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.727702Z digest=sha256:93c58e993b6452d48b6a39a4be4151f5386b32e98503540441c2416e5bc9b301

Observation 0a677e7e-f101-400f-b5ad-a44a2a77e18c · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.730593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.730593Z digest=sha256:9d52b942aa6b5755a0e57f792611368539e5cf1679ac585015ca0d02d6d614f0

Observation f6822db5-7cbd-4ebd-8023-6aa5e8df054d · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.734088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.734088Z digest=sha256:57256741b0648a245dc62dde1e3b53dcad19f89bac4c7b50eb8469ef808449be

Observation a3bac1f9-6d6d-4692-8fc9-d665278ec2a2 · outbound

This paper cites Small Language Model Meets with Reinforced Vision Vocabulary.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Small Language Model Meets with Reinforced Vision Vocabulary

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.737431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.737431Z digest=sha256:1464ddccaf982ccf216c4992b819a203109edf326a32e71b6aebd8424c2074a5

Observation badc2f20-4a8f-473c-a9a0-017195bb0e1f · outbound

This paper cites Mamballie: Implicit retinex-aware low light enhancement with global-then-local state space.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Mamballie: Implicit retinex-aware low light enhancement with global-then-local state space

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.058338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.740286Z digest=sha256:0ad0e95370176d8f19a6cbb4efde8968f989a5423c7279b81d6a9e4649cc0690

Observation 134df1f3-fcb4-491a-adcc-9379af1190e9 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Next-qa: Next phase of question-answering to explaining temporal actions

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.049933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.742620Z digest=sha256:efe62e49eabb06ee9bb6f24d8c3078686da5029daf98d7347849a8f6b1406675

Observation d607b380-1909-462d-9095-53da1f0b70c4 · outbound

This paper cites Robustmq: benchmarking robustness of quantized models.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Robustmq: benchmarking robustness of quantized models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.040207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.745037Z digest=sha256:9cc4532d2c7f8bf19ef5317bc8a507a6f64fc8be818cd522cac0fe41d7c4a20d

Observation be08aeb6-a7ee-4f92-bbee-39f6c09fb38e · outbound

This paper cites Novel object synthesis via adaptive text-image harmony.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Novel object synthesis via adaptive text-image harmony

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.030420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.747558Z digest=sha256:2810dd87d89243c74e9a39a9252cedb8116e243b5528dbafe7e20043e6f1bbcc

Observation a3ae0252-04b4-478c-b369-08b8ebd2b89b · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.749972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.749972Z digest=sha256:353f118d5b8046d54c7bbe17d4a008b4d65fd85f43451c087e7e19265713726d

Observation ae70fbcb-4bca-4025-b71c-330e23fd0387 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.020310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.752672Z digest=sha256:6742d8e38dcbbe6e016976be2e0a4d3c032db54b4922aa18b765bfb54c447d13

Observation fef99eab-0e7f-4355-80be-4bc14eb44e66 · outbound

This paper cites TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.755161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.755161Z digest=sha256:bb9e1555387ee07214c9ee6f48bdd5bfd14e03c2712be3a7a0ba4e3e327472f7

Observation 4d754c0a-2b72-492c-9eec-8d3682e7244c · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.757857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.757857Z digest=sha256:f870830eb6a11f1821b7334f7243f3cdef0ec01d8e856ff609788f352b5222c4

Observation 1c2d8b2c-e692-45ed-96e0-d939d4fe532c · outbound

This paper cites Lmms- eval: Reality check on the evaluation of large multimodal models, 2024.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Lmms- eval: Reality check on the evaluation of large multimodal models, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.761430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.761430Z digest=sha256:7611fd928618e01fc77742a5f481d9b6b2dbd6c92c4228297239cc2085e4ad4a

Observation 096daa0b-aece-40b1-b582-65cb437d3be2 · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.763783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.763783Z digest=sha256:28600102b6789f9936a53806b0aba634d8218c2314b72c443d526f0d58fbb137

Observation cc4c9fcd-7491-4696-ace4-b1513a46cc7c · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.766777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.766777Z digest=sha256:afb8365d82f859b7eeb3412369fa7aaf1ba80af4aed0103da8fcd699cbf5c2ad

Observation 1056b463-59c5-4538-8142-48d35f621a2f · outbound

This paper cites Llava-phi: Efficient multi-modal assistant with small language model.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Llava-phi: Efficient multi-modal assistant with small language model

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:24.007353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.769618Z digest=sha256:d603a8b9afbe39ffc80264fab950534ba6b1e9a35ee868af01889fd2045e122d

Observation 542b4934-4004-4159-bff7-cf40f1ae2ebb · outbound

This paper cites counterfactual reasoning.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models counterfactual reasoning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:40:23.998181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:40:23.771906Z digest=sha256:f004810294b36ac77febeed992189a687fcaaf5b00110642f7c8ab9d5edf9c8f

Pith citing papers

Observation 07f9dd78-b7ea-4ce0-afbc-03b2027df9ca · inbound

LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval cites this paper.

LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:31:40.742289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T14:26:59.015559Z digest=sha256:a92c015fa0f17ce0bacc591c8f8c44e39eeedea88fa19e472a4bd603c903b2a7

Observation 3f1acf08-3f3e-4d4d-bf9c-889c2570e0ce · inbound

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models cites this paper.

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:09.432264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:08:09.432264Z digest=sha256:eda1f97d9b32f1ba1678e2a148111d867d56cd826bb44eecfcbd11a4ea261240

Observation f2792e61-63a5-472a-8d51-e05cbbedfa04 · inbound

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding cites this paper.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.362186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.362186Z digest=sha256:7ee0a23fc5b3d72dbc6ac8c0766af65b2a5ac0fd3ff8d1732833b15e01fd369f

Observation 23f69557-35b7-43b7-9b7c-2d0d26d6a526 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:32.266783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:32.266783Z digest=sha256:5b7adbb58cab74036f1d3405a035a75bf55dcf0fdd792cd11fe8cd9dbb7fbb91

Observation 9fdda23a-e379-4a88-8fda-78e5a5fba147 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.595717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.595717Z digest=sha256:50324b0f91fe3381ae7e8aae6e59674465179d405302010f3e9b6ceb1a809c96

Observation 09cc8165-9107-402e-b976-e1c701189c49 · inbound

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs cites this paper.

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.139852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-19T03:18:11.993413Z digest=sha256:de4c314d9c7b20fedd357de6162fc270e4a25fd82d468418bf652469600a256b

Observation 2b08a16f-941a-4a64-8f19-744306d50b72 · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:53.927169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:88a182ded05a462970192a206a9379081f2b81029881c81034a26b29f308698e

Observation 6c70ceb7-65d6-46d6-a34a-39eabd1dee2e · inbound

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding cites this paper.

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:31:01.468199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T14:50:37.022338Z digest=sha256:64645b9b44daa7245dedcd9abde6b0acb59fd9ac34e21e73ec626a457e4d5104

Observation e3d60966-5cf6-41f8-a06f-eb0f69c2434f · inbound

TTF: Temporal Token Fusion for Efficient Video-Language Model cites this paper.

TTF: Temporal Token Fusion for Efficient Video-Language Model DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:06:00.550627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-11T01:55:32.319344Z digest=sha256:b418cfa21576a7a4a8f77bc9170319bb940c7df5406255b603150ebd04b1790f

Observation 2dad76de-74ca-4d25-993a-6a3554dad9ed · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.549583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.549583Z digest=sha256:b8323392c0df333d62a20a652f475ce5fd99aa5b6ef857b01f5d0b2287f5c0de