Pith. sign in

Paper Citation Record · LEDGER

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM

As of 15 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 4 inbound Pith citation observations for arXiv:2412.09530.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09530 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:02:17.077169Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:55:50.128159Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T00:11:23.194836Z

Reference resolution

82 of 82 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c932856-14c5-4282-a93b-74573a9ff5c9 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.659618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.659618Z digest=sha256:e1a85e51782b04c50a833352ec2abd746fc0dcb99a89a71579edc7db708a8151

Observation 2f64b1e1-2985-46c9-b3a2-babaf8aa917f · outbound

This paper cites Vizwiz: nearly real-time answers to visual questions.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Vizwiz: nearly real-time answers to visual questions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.665029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.665029Z digest=sha256:44c13bff1c7bf782b3b7307e708778b03d138d064e7e1e52de73c70502df4ebf

Observation f475a995-f86e-49aa-9937-fb3a814dab84 · outbound

This paper cites Latr: Layout-aware transformer for scene-text vqa.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Latr: Layout-aware transformer for scene-text vqa

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.409082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.670089Z digest=sha256:89f2818dd5319d3e2a2dfcdf50807adaa3292429bb7df7c5e2e2a4bb4700c5d1

Observation 15fb44d8-4c6a-497c-9033-a99c8806a409 · outbound

This paper cites Token Merging: Your ViT But Faster.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Token Merging: Your ViT But Faster

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.674915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.674915Z digest=sha256:20c337e08ec2157003e577000c96248280c6e7b36aa46e678a5edc6ca7f7e255

Observation 9e13b171-8b88-4956-86db-e6afc2e445b7 · outbound

This paper cites Visually Dehallucinative Instruction Generation: Know What You Don't Know.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Visually Dehallucinative Instruction Generation: Know What You Don't Know

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.680001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.680001Z digest=sha256:6d560daec999e2a5732c91fcf7c8356b59ca4439b5ba4ee063e065767d6678c0

Observation bf5da2d5-0b0c-4272-b7e5-f9a7c4d331ff · outbound

This paper cites Collecting highly paral- lel data for paraphrase evaluation.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Collecting highly paral- lel data for paraphrase evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.684985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.684985Z digest=sha256:7ddb4f95c5b869687ba4a805c3c657b36230396156935cc31e81f70ecba73e3f

Observation 48213966-d132-4252-8de1-d519eccc8b1a · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.690020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.690020Z digest=sha256:5f2d24f91e235267d045e72b3d3a4ce7d4b4f4bfaabada18ea784fb6306852f7

Observation 4b22c483-92df-4f72-b2c7-fb677c6e6c7d · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.695016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.695016Z digest=sha256:5323da827a2a39e6948878a470606474d451e0cad63534d26fe88b1543fc4f5d

Observation f3e638ce-8e0b-49ae-820f-b5b3eae0dcbd · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.699990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.699990Z digest=sha256:0284cbff93c63574fd8c92d4af6fdfdcdab736a89ac120fb16256f723dcab781

Observation a2949096-c0f9-4c23-990b-a3306ac34c6f · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.705603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.705603Z digest=sha256:990ca37793a47845113a73d19656775f58b9dca16749d20b762f9878ad808dd7

Observation 490402d5-bcf3-43a2-97a4-b673d1fd8f1b · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.710527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.710527Z digest=sha256:78df7b09e7add8d2014c6b43e6056d337421c7f702057f5ae82d9104c6dd82a4

Observation fa732190-ca30-4a3b-ae21-d0b373554497 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Gonzalez, Ion Stoica, and Eric P

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.715277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.715277Z digest=sha256:9e6627b802d39fa7c8b1788a444a21d059d4a2d2535a9b135607c52bd550d1f6

Observation 83e09b32-9781-487d-8b1c-819d3edcc7ad · outbound

This paper cites Comprehensive multimodal anno- tations with gpt-4o.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Comprehensive multimodal anno- tations with gpt-4o

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.373663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.720030Z digest=sha256:fb1bfa91181e303c250dc8348f16ca8fd409cd4e098ec4bc3d70aea2b1646a8a

Observation e6320788-acfe-4476-b34a-e682bcd4355e · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.724569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.724569Z digest=sha256:8e2ad09f8817e21f57ef67bb535817aa7b9f8995749bb7385478f7c42d38fc41

Observation 9a4abe3a-1186-44f9-8017-9f07c09ae370 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.729226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.729226Z digest=sha256:734ce45909d0f18a2c9d0b923d358a7739686dffd1e74640d520debbd417dfea

Observation 40cd1b06-4f20-4076-bffb-f783a3c8295b · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.734331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.734331Z digest=sha256:4bef5d645f1bbece75a54be4ddcadbd763e3cfc0988dc636b280c478a3a46429

Observation ff7ece3d-d2e8-4f32-8076-23f0d54cff9f · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.739084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.739084Z digest=sha256:d5f3429e27e1563619b5899b940f0f9edb46f3e3f2198904ec78feac6a09255c

Observation cf100765-ca48-4d74-9e1f-1d3ab3158cda · outbound

This paper cites Ai2d-rst: A multimodal corpus of 1000 primary school science dia- grams.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Ai2d-rst: A multimodal corpus of 1000 primary school science dia- grams

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.346844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.744616Z digest=sha256:0b42c2bdc2491efa5970ebf2819d902112159b34e9da8f9f85893b60e2cb4378

Observation ffd6d08f-7dfe-43fe-ac0a-881da6e23eda · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.749776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.749776Z digest=sha256:4f2a22783224dc828c538f7cc341272616af09f619e7e7214e1418149fe04977

Observation 456f3aae-7677-4c2f-94be-377df605d927 · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.322013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.754337Z digest=sha256:f49bb6dcefa77e9bb251c0070e6b0d29e337c046297dfdf5d6c36f03cabc09e0

Observation 250a2ceb-c93c-4548-98a5-97ec93f39d40 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.758859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.758859Z digest=sha256:59263a5773be4c303fb8b19b5b64e7cde6d9003aee8f4035b0e2e6ea7d3ddb6a

Observation f7430251-2754-4022-9482-1dff7aab992a · outbound

This paper cites Dvqa: Understanding data visualizations via ques- tion answering.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Dvqa: Understanding data visualizations via ques- tion answering

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.294724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.763334Z digest=sha256:ffff034aad3b8b82d7e94e9d0081be3bf2bac6307f9476024ac51fedb6d460ae

Observation e1bbb654-f8fa-40d2-b884-17d0a9222ea3 · outbound

This paper cites An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.768583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.768583Z digest=sha256:7e1b1ebe27b56dc43c3636fe87861079a60c60af2ce39b4b460cf05ab5e7ac6e

Observation 2eaed54d-a3a4-4d10-8929-5eea7e52c59b · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.278886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.773526Z digest=sha256:716a9c808546e6ccc0d0c6881e8786dfd8fec175d0e89077dde9a1dd75a86ddf

Observation 41b845fc-c6aa-48e9-a2ca-6f19c90c22f1 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.778226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.778226Z digest=sha256:c0c0d1101921930f4d0ff3673302a802cc50d4c4f931066c748eb5aa816fc7e5

Observation a0bfaa25-9e65-4ef6-a074-7fc360f0205d · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.783660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.783660Z digest=sha256:db1ecfa0feb56158c45f4d56ff09e7f7c032ad240938e6957e01d5e5bbad0d10

Observation c5416b6b-4ea6-4668-b65d-574e553a1237 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM VideoChat: Chat-Centric Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.788592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.788592Z digest=sha256:8889ff4cf693a5f4e8c367e9c0b89a0d8c30d697246beee95e74f09907b2f5e1

Observation 5748e27a-48d6-4d3c-ba12-252e9c687df2 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.262355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.793851Z digest=sha256:e9234ca6e1a3b22b4c5388f039cf8a12a4196832cde927af9d92efd93f4e18ee

Observation a0b9bd6b-a7ca-4147-8116-e3f91a5bc32a · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.799432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.799432Z digest=sha256:cc5d1c9aee936fb22ec8ebb0259a83c991b207ba025b3fb991642b80863a95f1

Observation 8c7cbd6e-9714-4531-9133-d38be1457413 · outbound

This paper cites Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.804403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.804403Z digest=sha256:dd2b28c60b73b9f9857b751ef302106f787a567d2da827a9a6d2c7b7827ecac3

Observation 6dad16e9-be53-47eb-8cec-192c10a466bc · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Llama-vid: An image is worth 2 tokens in large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.236859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.809186Z digest=sha256:f44a323d4dd234dce86df024fe4e8da222132a712d7a119966fcb50b5924b65b

Observation ab6005c4-ca54-47ff-a040-0aa55b6d446d · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.813771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.813771Z digest=sha256:df54d27a24f7b2970b84005c5b54e333bff3a26f535e5d2c6255ce9796caea4e

Observation d45fcce9-3d64-40ed-8728-194a7e5191a6 · outbound

This paper cites Vila: On pre-training for visual language models, 2023.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Vila: On pre-training for visual language models, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.220159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.818523Z digest=sha256:5de7f99ca51b328001cf9d4dbf7501aeda338048e572fbc1de1fd1ff9cdc867f

Observation fac194e4-b0ea-4dc6-9bdc-5b23850ed249 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Improved Baselines with Visual Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.827756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.827756Z digest=sha256:e100a79906b6eb8766f7ff131c935fa1d15f138e6f227bfde1c85768fde83d6b

Observation 3eb28c35-cb1c-4a6f-80de-48fa0c1b97ac · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.832375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.832375Z digest=sha256:09f0f292f8d2b798068845a8ff3398192169b385f166f9a325fa6cfb4f7617aa

Observation 5fe818af-c160-4879-946e-57db2d2bb695 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.836999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.836999Z digest=sha256:ef3a194fc431e4808161f2fefb4da4f458a4495279365329036fbfad3d7596ef

Observation 284278c8-33ae-405d-9c1e-67558e45bf4a · outbound

This paper cites Bt-adapter: Video conversation is fea- sible without video instruction tuning.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Bt-adapter: Video conversation is fea- sible without video instruction tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.841918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.841918Z digest=sha256:7f2dc6c883647e1ec1e1f010dee1571f9f2ecc5191a50a240300685eec5f993a

Observation da71166c-0baf-436d-bbe1-4adc66e9863a · outbound

This paper cites St-llm: Large language models are effective tem- poral learners.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM St-llm: Large language models are effective tem- poral learners

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.182353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.846449Z digest=sha256:a961997e574d42e1ee8afae02c2cc55f70676132061861e1ccda14a8e2a25a06

Observation 34ffb606-99e2-4492-9d0a-a2b6571a96eb · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM TempCompass: Do Video LLMs Really Understand Videos?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.851019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.851019Z digest=sha256:1f880aa7764202e244d7b242ec920dba7bc62f75ccfdf7b754b68b8537e6bd5d

Observation 669b8b86-cdfa-41d3-857e-2376f99d5cfa · outbound

This paper cites IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.856097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.856097Z digest=sha256:cea5785973dbe148294af337adb435b26509978b06f4d188b47d4da387a90a0a

Observation f11a2fd3-218d-4c88-92cc-c18cb430c74f · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.861339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.861339Z digest=sha256:67735d2b3d612c0d87f4770bb8db1d3d18ff331b2a983a6c19b9ae0c02667f3e

Observation 9b7884a8-4fa2-4ede-b530-da2dbe3cd041 · outbound

This paper cites Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.867167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.867167Z digest=sha256:5394e7d84fde431cf1f20b2c2ff71c110e003ba9ff718bc47c743b0f9fe33a15

Observation 22f37e24-5803-4a71-a39c-75bd88c0df7f · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.872311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.872311Z digest=sha256:78bb2b6a833ea96460e94d42e164a4820375ae1100e7c4ac5e9db4835c2189c0

Observation d2c9c74f-7125-461e-ac17-9c740481819d · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.878490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.878490Z digest=sha256:aa0b986b3a720a3de6beb389d0469e682bbcc6cb427d6cf47a11c4157569591d

Observation 96333313-a5f6-42ae-a4ef-a28d7438356e · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.167233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.885449Z digest=sha256:39863d0e3f84952bc50062b49cd9be98697502688bd0e454aec33280ddc3da51

Observation 176e8378-edc7-4376-9210-5ff6b92a69fe · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.151212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.892028Z digest=sha256:173a6c1b59c41303dff158b2cbe5d9a3d3b6c35d1d242b625e1401b635f6681f

Observation f85f3142-29bf-4a8d-bb07-3cacb4d4733d · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.897816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.897816Z digest=sha256:45c5cc65889eeb1596b6962f69e63a9fc292884e14aba1267f4430f62f0a74ae

Observation ae7ae494-fc7d-4c37-a695-bb0dcc0c8fc3 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Docvqa: A dataset for vqa on document images

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.135707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.902705Z digest=sha256:054b2311b2b54f0eedc16cf1011808fd16df5ef1a634a95f200e8a71a95996aa

Observation d9858a14-6a6c-4206-ac59-ed07a4a26af5 · outbound

This paper cites MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.907766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.907766Z digest=sha256:b88949ecfd234d5e0b7bd9266b6e9199964d5e1c112410fe15543ef04f3ad128

Observation 5cba445f-7eaf-4ed4-bba3-4ba6cab9bc0f · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Ocr-vqa: Visual question answering by reading text in images

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.912706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.912706Z digest=sha256:817ffde9d9ae5736b45c46716d67dddcc276b89fdc461b7e62cbb574c73260a6

Observation 4fe30b6c-f980-4d89-a14b-61c74321a606 · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence,.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Gpt-4o mini: advancing cost-efficient intelligence,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.917245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.917245Z digest=sha256:2e494d471df6b130300226723a28aaa03d344295435054aa62906e67431b4487

Observation 24494c93-8f86-4a94-b673-906357aafb5d · outbound

This paper cites Hello gpt-4o.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Hello gpt-4o

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.099165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.921837Z digest=sha256:7c3d71cfa4a08cdd094856aa066c6de004cbcff45bcc6860255ac85ce6628afd

Observation cfba6fe3-3d00-4afb-879a-0523880b494e · outbound

This paper cites Gpt-4v(ision) system card.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Gpt-4v(ision) system card

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.082386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.926381Z digest=sha256:e6d8bb839b48cc87f83a2856ce6c02008b5eca575319a14010a337e9d97d858a

Observation 86cad320-875c-488a-b7a4-a0bbf861d5a6 · outbound

This paper cites Per- ception test: A diagnostic benchmark for multimodal video models.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Per- ception test: A diagnostic benchmark for multimodal video models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.066389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.931070Z digest=sha256:99b794fcd2a47944309ed2d2a2fd96673cee6403d4d5a1f9bf845802d400f944

Observation e34ea4ac-e6b3-4628-801e-4f474bc95ed1 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Learning transferable visual models from natural language supervi- sion

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.935638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.935638Z digest=sha256:3a0722a08d5114b59b1527f838855fe0626b5e606778e1cd850d6d1e3f2177e3

Observation 6d55920e-3174-4702-95bc-0c69e8a77d29 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM A-okvqa: A benchmark for visual question answering using world knowl- edge

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.038930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.940493Z digest=sha256:dcb91c8e8f38fe80402b487ff65f9fe90f3dd9e671c8ae8e98d108d2022ca138

Observation b7d0812a-0a14-4666-a2df-e18542e37b32 · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.945563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.945563Z digest=sha256:572b4833ade76f500d6da73fdbc1251ec2e056916ca61b62116b0d29c661feda

Observation d2762b7e-6cb1-41a5-a532-2dc1a1988f12 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.950422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.950422Z digest=sha256:badb1e886e95c85ff7c12ae3b6f20cd6a5f006e59aa4c96ba10efabbfe7fb893

Observation beae8b2d-63e2-48cb-ad8d-f4eba0e23529 · outbound

This paper cites Qwen2.5: A party of foundation models.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Qwen2.5: A party of foundation models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.023268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.955254Z digest=sha256:1565f98ea65a8375ccdc7c926744eca50a720fa6012677cd46f6681f12be9e08

Observation c4758b66-8ec5-491f-8dc5-aef5f7dce3ca · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM LLaMA: Open and Efficient Foundation Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.959729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.959729Z digest=sha256:9c33c1586eec82832a8a334866c05cafb91ae57bb7f9b01a7dbab2ffd5ee6103

Observation 48ecf25a-20b4-4308-afd3-0a788910d382 · outbound

This paper cites How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.964433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.964433Z digest=sha256:d2508b0098405af3c96f8252c885a130f9137c992892981c2d73c911618939b6

Observation fc707a54-e503-4f03-b92e-a6388e5fff18 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.969355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.969355Z digest=sha256:c717904517e35d11dcb0693cb682b30025a5277fa61ad4ecfbb9314965f2f33a

Observation 618d974c-6669-45f4-ac63-4a0c6bc98b77 · outbound

This paper cites Elysium: Exploring object-level perception in videos via mllm.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Elysium: Exploring object-level perception in videos via mllm

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:18.007247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.974415Z digest=sha256:e0e1921c1cbcdf38d5f31aa10d4861edfb87041d8c2675f0e59ccac55bc7f595

Observation 60cfe8a5-9937-4b5d-bc77-ddbc1041062f · outbound

This paper cites Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.979622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.979622Z digest=sha256:ee34b448bec2eada479f6e22c6246e7fdf9f9c7790987cd583b65aef01c43005

Observation 7dc90eb2-36a2-4d1b-b069-d5b2e2e544b9 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.984664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.984664Z digest=sha256:6ae27589b9b4bc3b227fb56d89d65dd0ae0f7afdb33c09545f2a7f7995b8a1a1

Observation adc8c382-9e94-4f56-85ec-ca0908eea582 · outbound

This paper cites Q-instruct: Improving low-level visual abilities for multi-modality foundation models.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Q-instruct: Improving low-level visual abilities for multi-modality foundation models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:17.989489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:16.989890Z digest=sha256:cc9843c395e9d4271a302cf1ebcea8ee9be36c193399221e6abbf184ec6e5973

Observation 73ae1d19-5b5a-43b0-a087-b1761313f326 · outbound

This paper cites FreeVA: Offline MLLM as Training-Free Video Assistant.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM FreeVA: Offline MLLM as Training-Free Video Assistant

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.994805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.994805Z digest=sha256:8b134113862a86962502857d472daadb298992118ebf18975fe632142150120a

Observation ba097623-f394-4d1d-b78a-5d56ba900150 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Next-qa: Next phase of question-answering to explaining temporal actions

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:17.973769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:17.000406Z digest=sha256:98e99db638ade866a3ed21f8970b3ff9889650a6aa24f27e3dcdc213454b7f6b

Observation 3044a007-2726-4ce9-8d5b-632948f1254f · outbound

This paper cites Video question answer- ing via gradually refined attention over appearance and mo- tion.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Video question answer- ing via gradually refined attention over appearance and mo- tion

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:17.005834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:17.005834Z digest=sha256:621744f3600bc94373dc0da08a5e2d93fbaec1b01ad0be8528a97e7da78db74e

Observation 87d0de7b-e66a-447e-bf0d-8fa8f8cb0c7b · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:17.010894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:17.010894Z digest=sha256:f7448bbac2d47ca2d1536698dae72e5a74ddf354d4e9812c30ddbf41be39a9b0

Observation 33b63ce7-df92-4764-8b71-74501644df19 · outbound

This paper cites Ad- vancing high-resolution video-language representation with large-scale video transcriptions.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Ad- vancing high-resolution video-language representation with large-scale video transcriptions

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:17.945614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:17.016504Z digest=sha256:48a520253714cc13334e15057617a647f83e197817124cba30cda05f207eb96c

Observation 4c4fdf62-eb77-46d6-9424-feda94ec6a98 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:17.928317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:17.021280Z digest=sha256:0124f1d3d82749fa9c912865e229af6a53c2579e49722b0ea976f764079ed9c2

Observation 2f6514e6-704e-4001-b46a-e8d3ca9252b6 · outbound

This paper cites Sigmoid loss for language image pre-training.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Sigmoid loss for language image pre-training

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:17.026523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:17.026523Z digest=sha256:364847d8556d4fb4493aa3e58526254ebfefadc2accf489035e04de619a05d32

Observation c19181e4-2406-466c-a473-2748807052e6 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:17.031254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:17.031254Z digest=sha256:a5b64ce00d2ca3161fd0109e3846d783086eaf1283f0458550b808ba2276cc55

Observation 790fb985-21c2-43ed-b0e6-963eca332aad · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:17.036237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:17.036237Z digest=sha256:8df27f7894f7bfb5f1985767d154442261db0fa2434765a6b1f9c7b35b2f5b61

Observation cd321899-ed2d-4950-a8d7-7c2f4d6367e1 · outbound

This paper cites Long Context Transfer from Language to Vision.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Long Context Transfer from Language to Vision

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:17.046273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:17.046273Z digest=sha256:c4f51625dd068dbdceac8f781103991237c8b254e901f05a3b3ecef4ecde1769

Observation e88e98d9-66a4-404a-858b-53b2bfbb3211 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:17.050919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:17.050919Z digest=sha256:aeedeacd814ea0cacb0abd7212b5ee66f046d8667269ff65e4c9984768803edf

Observation 64d8fbe5-251c-48ab-a7ee-1d47237aa9d8 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:17.056130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:17.056130Z digest=sha256:4d68e130d4525776578cdfd1d8511f552d3b14050246295d8a29203e31e174f0

Observation b040157e-9a33-4a45-a7e9-9789ff492ae9 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Llava- next: A strong zero-shot video understanding model, 2024

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:17.901206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:17.061375Z digest=sha256:0338e994f19f5e0117f05e339b772ab3cef2392cc111ec959cf036bf72043bb4

Observation e7f489b0-07a9-4e3e-b889-79c40955f131 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM MLVU: Benchmarking Multi-task Long Video Understanding

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:17.066508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:17.066508Z digest=sha256:1de1868e1cc525508c8b573568f9e4b973d7a8801dedf0c92e903cad2fc02d5e

Observation 5fcb3e66-8bff-4fcb-9776-a3c1bd75ce33 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:17.071951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:17.071951Z digest=sha256:03449ab9c854f94e0672308599a8a524384cdc990c729d32cdcf71088cc66d87

Observation 3c48134e-fed9-4550-8d3e-278cc7b834ec · outbound

This paper cites Visual7w: Grounded question answering in images.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Visual7w: Grounded question answering in images

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:17.884648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:02:17.077169Z digest=sha256:d3e29e08c7eb42d655be923fef9b398d06cb3f47d1d0de0506fb91c97768a2cc

Pith citing papers

Observation 559084cc-f064-4cc9-8f56-f56b9d8f608a · inbound

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding cites this paper.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.371800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.371800Z digest=sha256:3dd0048686229866256303457ec4d86830cb07a74b1ba57ccdb6b0f2fae1d80a

Observation 172834a1-4ff0-44bd-980d-4158be0247b0 · inbound

Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval cites this paper.

Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:11:23.198043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T00:10:01.596422Z digest=sha256:eeae0d03b7fa664cd9cd3c8da6a56f73bad609de558062e842d35d14cfcf8c12

Observation b4359393-9c7c-47fa-85a1-44bdbd4b6697 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.659904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.659904Z digest=sha256:735459a4ecdd97746dabbb98c97efdfd3c4d134edfeaf950b36cb95763f3e2fd

Observation 5449c859-83cf-4de4-87c8-6b5172039234 · inbound

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models cites this paper.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.128159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.128159Z digest=sha256:7c2add07fa80ca00ab845849343abe8c9d4347905bb4453f058179a3826d8179