Pith. sign in

Paper Citation Record · LEDGER

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding

As of 19 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 4 inbound Pith citation observations for arXiv:2506.03990.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03990 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:56:05.432793Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:51.641946Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:19:29.938511Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa61887c-8890-4ecc-bfdf-96752c60cec1 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.984503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.276195Z digest=sha256:12178f850049f20ad03f4bb471d556a3b9f08569009cfceb107b7f7a49b9a8cc

Observation dc61cc26-beb8-430d-85f6-82c0ca483762 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.974300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.280476Z digest=sha256:4f786888955f9c4ae66e1768d68bf3e8909f50b6fbf5e8b2afbda69b4014d1bb

Observation c79db32a-ddd9-407a-a895-d22bc85a8470 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.963475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.284562Z digest=sha256:32b5f1163cff2809d2c2e9d8865a602754603c993b951655f4f3a0ff30f0bd5e

Observation e9bc4e7f-5737-47b7-bc18-3eda5692741f · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.952522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.288049Z digest=sha256:911a85bca0b33a7c1f8d603646a0634ee9dd9d0a0cc4d414a009723b57fd625c

Observation 2a779ba7-a508-414c-a14a-7e710aa26f10 · outbound

This paper cites An Introduction to Vision-Language Modeling.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding An Introduction to Vision-Language Modeling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.291412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.291412Z digest=sha256:dd32e6f5899ddda28abe1d4b7fc124a897bc4c8f2a9faea2dcaea522c84528aa

Observation d4c6e30f-30cf-4d31-b3da-6649ec251dc0 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.295051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.295051Z digest=sha256:20629eb1ee014244f713a1e3e60e2a95d6b3bf6c41c8f50f4b9c5c3d6631da8f

Observation 33af4884-a719-4c17-889e-bf36fe9aa410 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.298607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.298607Z digest=sha256:359ba191fe1f08d84539150f3a2e13348bf87169c7759015d43fa61c23e98ce8

Observation 9dc5e770-71cc-4304-8aad-4f9b10a4b007 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.302047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.302047Z digest=sha256:a5be2658341f9c9d00f05c60a5bd3a1f376ec9c3a4e24515401b15ae5cd46242

Observation 26f4b637-9060-4488-80eb-9c2a838b3e8f · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.305472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.305472Z digest=sha256:a09fffb8aa11d5998d769942852597ae072c6499cfa0a07b68d7a8dff20aa2cd

Observation a916a4f9-d205-473b-b4f7-13df30d1a489 · outbound

This paper cites FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.308899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.308899Z digest=sha256:bb5327c792db3757d71aae040218f7f96f26a49c6325e4353e5adf1b805ffb3e

Observation d5b5039e-25b6-42c4-90b9-336ed569aca1 · outbound

This paper cites Matryoshka Query Transformer for Large Vision-Language Models.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Matryoshka Query Transformer for Large Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.312460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.312460Z digest=sha256:376a70b9d2b3ac8276e3d8e192bcb8ccffa9b3221d906fed3969fc9fd2d08d92

Observation e6f80ed6-3d9b-4dc6-b009-6f5c563b7436 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.941866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.315715Z digest=sha256:ce3b313026e911dd4445c93842460a24555664153fca15e72032ebc7a8929b33

Observation da39ac70-4056-459d-bb73-49d2f7a73fe0 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.318608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.318608Z digest=sha256:afde44eca734c2f3bd0e8de26db26ba02b77c3dbbf7a5d348980cd0eefaee81a

Observation 264fb417-cc39-45a4-a218-e19a0cc9fa8d · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.321429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.321429Z digest=sha256:7281d7ca107a017682097428301a5f6d79f4f74d2c65cdb5eacb0550297068df

Observation dfbdef21-7869-468d-8320-59bfb76e569e · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.324395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.324395Z digest=sha256:413eaaf8f4d6e2c781bc68250dee49ea5a853b0f79d609747812c17c171ddd99

Observation fb72792c-376e-40b7-80b9-4ed91c6044b2 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.327297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.327297Z digest=sha256:0f678cc2e0318db4c840e665c544486208cdfc5f51f4650995e8a3ec96f056dd

Observation 13319851-0f2b-40b7-8362-c5ffbd386569 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.918195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.330290Z digest=sha256:225a14670edc87f9b2a39033eaca7e50d792264e5c85c34c3b7fa3be42630251

Observation 81818563-8738-439b-a096-f401ce88c5b9 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.333184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.333184Z digest=sha256:83226f68f9fbf1ec4f9ab72d73801e6156650b569029a73eded5eec4fb31bde8

Observation 2ac41062-a81f-4c23-8768-e9838f583c2a · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.336628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.336628Z digest=sha256:994f21f6c338be7e7f831dfd561e27ed45c0fff6e87b369843b6b583d2bf3bdb

Observation bac04eb9-4d82-49b6-bc64-11e7635497c8 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.339561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.339561Z digest=sha256:66d5ee17cb068935113a73425c67cc96500ab222bd76773013a6e39f9ba8d038

Observation 8f293930-0977-4141-a852-9528cbacc60f · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.900748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.342903Z digest=sha256:1da315af25e91091d53f2d6761d03ac30294c29b239b14ca68b84bb5cd454291

Observation b21742c1-cb48-48f8-9e18-72ff9f469fec · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.890243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.345833Z digest=sha256:fe60f1d472297083ebf4d38b2eeb4caef213dd25baf30e4a59e8a6a7253433b2

Observation 20152ca6-858c-4b66-ad82-e400bf1af8e4 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.879074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.348695Z digest=sha256:58f64194d6377a03dd3a3b309f4d114d97379454cce1f350be91b841dece30b3

Observation 0d1ff616-b72e-42c1-8406-557aaa997bec · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.868075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.351703Z digest=sha256:bfa03c7dde5021ebcc49228a3dd141834b3c1858069e8412bdfa2aa947369aae

Observation 53677ca5-1ac8-42b3-9335-e3510cf0ee6d · outbound

This paper cites xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.354802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.354802Z digest=sha256:64ec023960feebab6bcb816d6c48c0eba28943b32568e58720c804f91b37e4de

Observation 64ec724f-b4c6-4fcc-a125-e0bdab7eedd9 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.358235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.358235Z digest=sha256:c2e820de7dd9873f00e05cc048de9a7ddef43ccc30e15b43a319105e848c2b5a

Observation f2792e61-63a5-472a-8d51-e05cbbedfa04 · outbound

This paper cites DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.362186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.362186Z digest=sha256:c2dc52f3722323c2150d850abc20ef92411bfb9833296ee87fc4c773ab78a422

Observation 106473ca-fc15-413e-9740-5656892a4c6f · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.365504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.365504Z digest=sha256:2edb12723e7de8342a99f533de02351b5582c869dd1ad6b538cbb142f902d78a

Observation 64346559-2710-49cb-890b-c1fc11bfaf92 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.368592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.368592Z digest=sha256:201db76e92bbfa8b0053c029a9ead2ff44ae5ad919e5152d97ef618c41811824

Observation 559084cc-f064-4cc9-8f56-f56b9d8f608a · outbound

This paper cites Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.371800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.371800Z digest=sha256:767d843f901396929f6e3245161d120e31665ec914dd0c9622c031023ff67872

Observation 73c33537-a5e3-4dc3-b645-08885b5d2b8c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.374864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.374864Z digest=sha256:7fb14d6fc03eeedcfb3b508483bcc44515eda32d768cb04c8527efe9fadfc1e0

Observation 5f45281d-0bfb-4c17-81bc-01fc47dc6ea7 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.857436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.377969Z digest=sha256:209a61b57cbdab2ea6944b05ae51b2d3cbcf998d915baafcbac03660969b45e9

Observation fe533a48-10ce-4d8c-beb2-38ac03429af0 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.847252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.380731Z digest=sha256:69bfd2188407812c006e5688a9362f80d13c4e6926b4b68263fd62b9079208e2

Observation ca11976d-f9ad-4707-a57a-bbee01e75148 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.383591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.383591Z digest=sha256:4aa581de888b125acec704c345674faf768396de96169d89a6bbbcac6da2ca19

Observation 8ae79a53-3939-4976-a6e4-916cb7ee8073 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.835679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.386907Z digest=sha256:006b675cb37360178e8478e5bb42c28be6d016875dc67eb1067ff3d2384f8760

Observation b80eb3c8-54d4-45d1-975d-01eed4caa809 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.389890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.389890Z digest=sha256:d6d49d2497c7526ad99f72ee8cb6b6af2ea4e53675e254bc2c35e8cfc90fcaaf

Observation 6a589e1a-66ac-416d-80f8-67ffa2397b8b · outbound

This paper cites Qwen2.5 Technical Report.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Qwen2.5 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.392924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.392924Z digest=sha256:14a73f3e064a48b6fe683d8c518b2859811beaa7598090bf598657bf4ed3d44f

Observation c81d325b-2bb5-4b51-8784-28bdf7bd4a62 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.396090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.396090Z digest=sha256:12ff31a6e8de340d20c6028b9005ff1415c91dc5034b898aab0bbc41d2c2ba2e

Observation 2d1ab832-3f2b-4656-a918-af0566f5aa48 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.399011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.399011Z digest=sha256:d529a45fc082b62a96bd1d1a8520caab4c5fd1f70ab9c0aafd5ca9543b0aeaaa

Observation 6e6a6b50-4c14-491e-9a18-6c60e07f2cc0 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Sigmoid Loss for Language Image Pre-Training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.402373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.402373Z digest=sha256:6024eed7165f289970314334e6fb50e5aecf8e7ed52bf3c523ded656692fba1b

Observation 29db3ad9-d287-4e59-8754-7cbb96c444d0 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.407052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.407052Z digest=sha256:fa01fb162f1ae5d1a9a9d5c8ee3dd32860c77251cbabe3325f40d08f780995fa

Observation d1f0c0cb-268c-402c-aad2-c20d039f92d4 · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.411061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.411061Z digest=sha256:e3274508c849d6936c37e23d9013e2221079bf57e922a6dbdfc9c40417fe576b

Observation 500a1ece-fc59-4481-a16d-988122648280 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.418612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.418612Z digest=sha256:4776e38bfaa28fb03870b9591571571237ddb4237b36bdabc774e40aa133d9f4

Observation a0b9f381-ede6-48c1-903c-f0c0bcad59ca · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.421882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.421882Z digest=sha256:94538dd8612433e589e85a5d6508b0ba00006dee3b4074b92bf6eb6e3eb2220f

Observation 1ad1be84-66b9-4d36-9bc9-d868946137e9 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.425465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.425465Z digest=sha256:28de673dce6519d777dc742ae60c6abaf2a1e752bc723bedc297c7ef1c094806

Observation 4cc700f1-1743-4f53-81f9-3d7a5be55d0b · outbound

This paper cites online" 'onlinestring :=.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding online" 'onlinestring :=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.428727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.428727Z digest=sha256:97587e997bb2fcc9bcd360228575e9291de59ddde4bd99d7f568f07e4fa89baa

Observation b331cc54-46dd-4cba-be33-02492310fb91 · outbound

This paper cites write newline.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding write newline

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.432793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.432793Z digest=sha256:3da25717dd71621738d9f18b143fc181c44512bec5ddd6cb59a945dbe0508284

Pith citing papers

Observation 5b063c2e-aa4d-4da7-973c-621e23709d06 · inbound

Stateful Token Reduction for Long-Video Hybrid VLMs cites this paper.

Stateful Token Reduction for Long-Video Hybrid VLMs DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:05.052870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:05.052870Z digest=sha256:66fc4caf50320ea2a9c0fc7e60b62755635fe70219d4624204ee250facd90d7f

Observation 3bcde7a0-8149-44a2-85f6-c3ce2e762161 · inbound

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference cites this paper.

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:19:29.940719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-26T18:17:53.013043Z digest=sha256:e775b432c7a8c9d34a35956cb7a32f0b5e12674f9c84b64b80b75d3259a18766

Observation 62bd9af1-87b9-4723-8fd0-00f2d8d2e4a5 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding

Reference 248

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:55:42.206405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:e64c4bff7488bd025d95b563f2080379007bcda2eb6e9a5ad6cfa775c429bab3

Observation 84d2da63-359b-4e4c-9943-42b6095c3fb3 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.641946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.641946Z digest=sha256:d979de032c8abafec351e57215ed4bcb6814c24263c6e8b16536bf1b19177879