Pith. sign in

Paper Citation Record · LEDGER

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding

As of 9 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 4 inbound Pith citation observations for arXiv:2506.03990.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03990 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:56:05.432793Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:51.641946Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:19:29.938511Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa61887c-8890-4ecc-bfdf-96752c60cec1 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.984503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.276195Z digest=sha256:bdea410644bf141a2ade739500579ed28c4cdd2f0f159fe84411cebd1b3a18b3

Observation dc61cc26-beb8-430d-85f6-82c0ca483762 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.974300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.280476Z digest=sha256:534b4b9d7db3ea9e583b9f0a56d92467704979740df349ac83d90c3e0994c84e

Observation c79db32a-ddd9-407a-a895-d22bc85a8470 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.963475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.284562Z digest=sha256:1c557899bf6f6c075469fa4b8ce5d7e92c86fd3e3077c225d2ac485c831f6c4a

Observation e9bc4e7f-5737-47b7-bc18-3eda5692741f · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.952522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.288049Z digest=sha256:4ac5a40aabb9b1caaf56153d33a5ed35ce3143d2f49cad06c3bfae62ea7bfa83

Observation 2a779ba7-a508-414c-a14a-7e710aa26f10 · outbound

This paper cites An Introduction to Vision-Language Modeling.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding An Introduction to Vision-Language Modeling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.291412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.291412Z digest=sha256:276d909656a0b2ed0abe2e11be359ea2e3b5e00b68a5ebe976790a2e2844b375

Observation d4c6e30f-30cf-4d31-b3da-6649ec251dc0 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.295051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.295051Z digest=sha256:af7e99c962ec65b705824afce4341e156e3edd73c308965cf554bff16c9a916d

Observation 33af4884-a719-4c17-889e-bf36fe9aa410 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.298607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.298607Z digest=sha256:0acfab7e8008e7fd577c05711cfbb1a5901d0eb4cb846a47515027ac4fa6b83f

Observation 9dc5e770-71cc-4304-8aad-4f9b10a4b007 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.302047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.302047Z digest=sha256:7b9597ee1c24fc7cdf10a1c1aeb6d85b6767895734c940913271239d20924091

Observation 26f4b637-9060-4488-80eb-9c2a838b3e8f · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.305472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.305472Z digest=sha256:6a3bd0cd7483ae909187260b4f95bc49dafd033df009cfe3f3f46351b3d8a6f4

Observation a916a4f9-d205-473b-b4f7-13df30d1a489 · outbound

This paper cites FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.308899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.308899Z digest=sha256:cf40b5e68f95ea5d673598b95b34f1e27420f838266d718a04512d910d0c2812

Observation d5b5039e-25b6-42c4-90b9-336ed569aca1 · outbound

This paper cites Matryoshka Query Transformer for Large Vision-Language Models.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Matryoshka Query Transformer for Large Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.312460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.312460Z digest=sha256:dc003b5e612f2f6be8d1e8a5230d2ca48dec9d0b6cf57ff46e65f404595ff337

Observation e6f80ed6-3d9b-4dc6-b009-6f5c563b7436 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.941866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.315715Z digest=sha256:fa634a39b0b49e27cf3ee7bce600634d96f833fb7550af4112b62b89b4fef0ee

Observation da39ac70-4056-459d-bb73-49d2f7a73fe0 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.318608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.318608Z digest=sha256:1a99d2186e7055c4c35fdb0cbad4a1598f30c57c9452eb787e209ccb04031656

Observation 264fb417-cc39-45a4-a218-e19a0cc9fa8d · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.321429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.321429Z digest=sha256:90a6b227c77d2cef32fbacd36e3d7acbf41e98b42ac2e22f1b982f36f52feccf

Observation dfbdef21-7869-468d-8320-59bfb76e569e · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.324395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.324395Z digest=sha256:70c3efc22e86ef326345794d7ed15325e2202f1e0532ed80a4ed7d27e5cdfff4

Observation fb72792c-376e-40b7-80b9-4ed91c6044b2 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.327297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.327297Z digest=sha256:55cacea76856210456ada2603db268d1f12c4cdc3b86bc3cd2dc224da5a0f873

Observation 13319851-0f2b-40b7-8362-c5ffbd386569 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.918195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.330290Z digest=sha256:d7f562fd332ce1054369c454a8d68c700cdaa210fe0a316fdf76c4c366af6358

Observation 81818563-8738-439b-a096-f401ce88c5b9 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.333184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.333184Z digest=sha256:e2c589669f8a0930b8b5ffd2615bc5f4ba09cca4c964ae6e4f8164a6c584c33b

Observation 2ac41062-a81f-4c23-8768-e9838f583c2a · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.336628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.336628Z digest=sha256:3ffda06d687187fa389e34e3c482a9c0e62659e9755258abd4e80fa9aa3bd1e8

Observation bac04eb9-4d82-49b6-bc64-11e7635497c8 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.339561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.339561Z digest=sha256:204b84f1a8052e3f612da105958acf0e6a43ce9a737617df9a46f1127237a80e

Observation 8f293930-0977-4141-a852-9528cbacc60f · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.900748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.342903Z digest=sha256:4fdd0471e7a985a5b68884b486995de7c1cc6ee3601622915aec18a5298d6af4

Observation b21742c1-cb48-48f8-9e18-72ff9f469fec · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.890243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.345833Z digest=sha256:a492a5d56e855362217e2de86aca9c51014c7262a6338a4ec7afd3675dc94008

Observation 20152ca6-858c-4b66-ad82-e400bf1af8e4 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.879074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.348695Z digest=sha256:590fef7067d4ec14d1bd1d6073c1cc0c999167875c2cb24fced5b3d1cfc1a448

Observation 0d1ff616-b72e-42c1-8406-557aaa997bec · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.868075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.351703Z digest=sha256:53ddc83e354dc6fa56ba0ea53bff19afc64a138ef801f226dfdca7fba3e14d8c

Observation 53677ca5-1ac8-42b3-9335-e3510cf0ee6d · outbound

This paper cites xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.354802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.354802Z digest=sha256:ef28bb80c8fa29316a255552cea397d758a9d25a75c5fcf03576ae285775db7a

Observation 64ec724f-b4c6-4fcc-a125-e0bdab7eedd9 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.358235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.358235Z digest=sha256:961a67541b070483a28263c48b49365d0febd57db4a2810cbb49fe1c7cde0e84

Observation f2792e61-63a5-472a-8d51-e05cbbedfa04 · outbound

This paper cites DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.362186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.362186Z digest=sha256:734315fadec43082104b26e64cb3d83454682dd788692a4e499216e39be00f1f

Observation 106473ca-fc15-413e-9740-5656892a4c6f · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.365504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.365504Z digest=sha256:f249e1f702c0b49d8d0addd0a7aec5d12ad6f86d3e06edcade7af288b28216e7

Observation 64346559-2710-49cb-890b-c1fc11bfaf92 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.368592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.368592Z digest=sha256:1c740268ba8e717d49ba140a2624e528748b766ca80945095bef9ec686765862

Observation 559084cc-f064-4cc9-8f56-f56b9d8f608a · outbound

This paper cites Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.371800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.371800Z digest=sha256:b6ff7d208b12fdbaa1ac377c86afa6bf7511ee07c84d9370ac6959cfb12c0112

Observation 73c33537-a5e3-4dc3-b645-08885b5d2b8c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.374864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.374864Z digest=sha256:c0680da65297d9fddca12832b741e7f2a692298656f9a3af2cee5cee8545f46e

Observation 5f45281d-0bfb-4c17-81bc-01fc47dc6ea7 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.857436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.377969Z digest=sha256:2d6a3403222a91260771e65da6d18c25a04736290a0f083d75d413e10503395b

Observation fe533a48-10ce-4d8c-beb2-38ac03429af0 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.847252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.380731Z digest=sha256:c170ae466330aa1a9ec50766eae36add2fb55927151b5ba63c186dc3f934008f

Observation ca11976d-f9ad-4707-a57a-bbee01e75148 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.383591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.383591Z digest=sha256:8ad666023ffa908097a24e28ce804fd4f9af166523641db3025339f574121105

Observation 8ae79a53-3939-4976-a6e4-916cb7ee8073 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:56:05.835679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:56:05.386907Z digest=sha256:df0c7c46b12547011fe161a10fc89fa83608ce8b07f850126bb412efc488c01f

Observation b80eb3c8-54d4-45d1-975d-01eed4caa809 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.389890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.389890Z digest=sha256:55129ced2d3d3a82b525d053c1bdc980581c8a8e42705063185344792c9f16d4

Observation 6a589e1a-66ac-416d-80f8-67ffa2397b8b · outbound

This paper cites Qwen2.5 Technical Report.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Qwen2.5 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.392924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.392924Z digest=sha256:cc97b2e7f9fde061de6aaccccdfb0f7aec0536bfdfd2f8bed26d499edbbf994d

Observation c81d325b-2bb5-4b51-8784-28bdf7bd4a62 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.396090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.396090Z digest=sha256:5ad6dcd09d8bec551d0ad24d64fccdb624f79020ad6c640cdc7107ab15a1f5bf

Observation 2d1ab832-3f2b-4656-a918-af0566f5aa48 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.399011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.399011Z digest=sha256:ef06425f4b072a3fa773533e5fb026cb3a159c6dfdb9b504f71a4872f62b386d

Observation 6e6a6b50-4c14-491e-9a18-6c60e07f2cc0 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Sigmoid Loss for Language Image Pre-Training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.402373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.402373Z digest=sha256:86e55f2a8fd18808756b6ccf1efdd541f20a1865bd9254d7ff0771017999f0f9

Observation 29db3ad9-d287-4e59-8754-7cbb96c444d0 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.407052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.407052Z digest=sha256:ac72b8622d55f1401fd40c03d677cf7e64c91ac3f0e011029642bb3e00cd683d

Observation d1f0c0cb-268c-402c-aad2-c20d039f92d4 · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.411061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.411061Z digest=sha256:028577d848472d5f9829a1d7340038ead772372034601a5dd5d226fd76858cf3

Observation 500a1ece-fc59-4481-a16d-988122648280 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.418612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.418612Z digest=sha256:025797d4ba00592c0d2b7288fba61d87d31c5613166755c241c7ee4b088e623d

Observation a0b9f381-ede6-48c1-903c-f0c0bcad59ca · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.421882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.421882Z digest=sha256:a60f880aa2f38f15182f10cdb8230156934b9404a5833ad979a184cdf8702ba9

Observation 1ad1be84-66b9-4d36-9bc9-d868946137e9 · outbound

This paper cites an unresolved cited work.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.425465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.425465Z digest=sha256:a1b0d26ab5c28389278b1e69011f34787a610a16f8fa05f1968b5f3251440fb2

Observation 4cc700f1-1743-4f53-81f9-3d7a5be55d0b · outbound

This paper cites online" 'onlinestring :=.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding online" 'onlinestring :=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.428727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.428727Z digest=sha256:79c3751fb079957ff184911693bbae25e9f33dcf4086ea1179a41616a39cd568

Observation b331cc54-46dd-4cba-be33-02492310fb91 · outbound

This paper cites write newline.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding write newline

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.432793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.432793Z digest=sha256:01976dbbfcb55d21ec679b3942e8b38252b878f58c92d368652a385e93be5b4e

Pith citing papers

Observation 5b063c2e-aa4d-4da7-973c-621e23709d06 · inbound

Stateful Token Reduction for Long-Video Hybrid VLMs cites this paper.

Stateful Token Reduction for Long-Video Hybrid VLMs DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:05.052870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:05.052870Z digest=sha256:5143316f36ba8b9ce309d0e0c8aaaf048eac27d878ab85be5a121e80021c983d

Observation 3bcde7a0-8149-44a2-85f6-c3ce2e762161 · inbound

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference cites this paper.

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:19:29.940719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T18:17:53.013043Z digest=sha256:ee2c8e9bc68529979131460521c8cd0345b357d442ce4a7f53ad7087f5522398

Observation 62bd9af1-87b9-4723-8fd0-00f2d8d2e4a5 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding

Reference 248

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:55:42.206405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:da49ac493dd2eb2539c009e9795d45a084d12130ef5ac17c89d99db1b43c627a

Observation 84d2da63-359b-4e4c-9943-42b6095c3fb3 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.641946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.641946Z digest=sha256:bf2f113dec47d5ffb3a80e3d2991b58a4b133b370abb0f3ab9834ea147214e45