Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:02:17.077169Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 4 inbound Pith citation observations for arXiv:2412.09530.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:02:17.077169Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T00:55:50.128159Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-17T00:11:23.194836Z
82 of 82 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5c932856-14c5-4282-a93b-74573a9ff5c9 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f64b1e1-2985-46c9-b3a2-babaf8aa917f · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Vizwiz: nearly real-time answers to visual questions
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f475a995-f86e-49aa-9937-fb3a814dab84 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Latr: Layout-aware transformer for scene-text vqa
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 15fb44d8-4c6a-497c-9033-a99c8806a409 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Token Merging: Your ViT But Faster
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e13b171-8b88-4956-86db-e6afc2e445b7 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Visually Dehallucinative Instruction Generation: Know What You Don't Know
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf5da2d5-0b0c-4272-b7e5-f9a7c4d331ff · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Collecting highly paral- lel data for paraphrase evaluation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48213966-d132-4252-8de1-d519eccc8b1a · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b22c483-92df-4f72-b2c7-fb677c6e6c7d · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3e638ce-8e0b-49ae-820f-b5b3eae0dcbd · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2949096-c0f9-4c23-990b-a3306ac34c6f · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 490402d5-bcf3-43a2-97a4-b673d1fd8f1b · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa732190-ca30-4a3b-ae21-d0b373554497 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Gonzalez, Ion Stoica, and Eric P
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83e09b32-9781-487d-8b1c-819d3edcc7ad · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Comprehensive multimodal anno- tations with gpt-4o
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e6320788-acfe-4476-b34a-e682bcd4355e · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a4abe3a-1186-44f9-8017-9f07c09ae370 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40cd1b06-4f20-4076-bffb-f783a3c8295b · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff7ece3d-d2e8-4f32-8076-23f0d54cff9f · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf100765-ca48-4d74-9e1f-1d3ab3158cda · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Ai2d-rst: A multimodal corpus of 1000 primary school science dia- grams
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ffd6d08f-7dfe-43fe-ac0a-881da6e23eda · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 456f3aae-7677-4c2f-94be-377df605d927 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 250a2ceb-c93c-4548-98a5-97ec93f39d40 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7430251-2754-4022-9482-1dff7aab992a · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Dvqa: Understanding data visualizations via ques- tion answering
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e1bbb654-f8fa-40d2-b884-17d0a9222ea3 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eaed54d-a3a4-4d10-8929-5eea7e52c59b · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 41b845fc-c6aa-48e9-a2ca-6f19c90c22f1 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM LLaVA-OneVision: Easy Visual Task Transfer
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0bfaa25-9e65-4ef6-a074-7fc360f0205d · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5416b6b-4ea6-4668-b65d-574e553a1237 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM VideoChat: Chat-Centric Video Understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5748e27a-48d6-4d3c-ba12-252e9c687df2 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a0b9bd6b-a7ca-4147-8116-e3f91a5bc32a · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c7cbd6e-9714-4531-9133-d38be1457413 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dad16e9-be53-47eb-8cec-192c10a466bc · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Llama-vid: An image is worth 2 tokens in large language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ab6005c4-ca54-47ff-a040-0aa55b6d446d · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d45fcce9-3d64-40ed-8728-194a7e5191a6 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Vila: On pre-training for visual language models, 2023
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fac194e4-b0ea-4dc6-9bdc-5b23850ed249 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Improved Baselines with Visual Instruction Tuning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eb28c35-cb1c-4a6f-80de-48fa0c1b97ac · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fe818af-c160-4879-946e-57db2d2bb695 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 284278c8-33ae-405d-9c1e-67558e45bf4a · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Bt-adapter: Video conversation is fea- sible without video instruction tuning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da71166c-0baf-436d-bbe1-4adc66e9863a · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM St-llm: Large language models are effective tem- poral learners
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 34ffb606-99e2-4492-9d0a-a2b6571a96eb · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM TempCompass: Do Video LLMs Really Understand Videos?
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 669b8b86-cdfa-41d3-857e-2376f99d5cfa · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f11a2fd3-218d-4c88-92cc-c18cb430c74f · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b7884a8-4fa2-4ede-b530-da2dbe3cd041 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22f37e24-5803-4a71-a39c-75bd88c0df7f · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2c9c74f-7125-461e-ac17-9c740481819d · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96333313-a5f6-42ae-a4ef-a28d7438356e · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 176e8378-edc7-4376-9210-5ff6b92a69fe · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f85f3142-29bf-4a8d-bb07-3cacb4d4733d · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae7ae494-fc7d-4c37-a695-bb0dcc0c8fc3 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Docvqa: A dataset for vqa on document images
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d9858a14-6a6c-4206-ac59-ed07a4a26af5 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cba445f-7eaf-4ed4-bba3-4ba6cab9bc0f · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Ocr-vqa: Visual question answering by reading text in images
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fe30b6c-f980-4d89-a14b-61c74321a606 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Gpt-4o mini: advancing cost-efficient intelligence,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24494c93-8f86-4a94-b673-906357aafb5d · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Hello gpt-4o
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cfba6fe3-3d00-4afb-879a-0523880b494e · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Gpt-4v(ision) system card
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 86cad320-875c-488a-b7a4-a0bbf861d5a6 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Per- ception test: A diagnostic benchmark for multimodal video models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e34ea4ac-e6b3-4628-801e-4f474bc95ed1 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Learning transferable visual models from natural language supervi- sion
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d55920e-3174-4702-95bc-0c69e8a77d29 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM A-okvqa: A benchmark for visual question answering using world knowl- edge
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b7d0812a-0a14-4666-a2df-e18542e37b32 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2762b7e-6cb1-41a5-a532-2dc1a1988f12 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beae8b2d-63e2-48cb-ad8d-f4eba0e23529 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Qwen2.5: A party of foundation models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c4758b66-8ec5-491f-8dc5-aef5f7dce3ca · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM LLaMA: Open and Efficient Foundation Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48ecf25a-20b4-4308-afd3-0a788910d382 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc707a54-e503-4f03-b92e-a6388e5fff18 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 618d974c-6669-45f4-ac63-4a0c6bc98b77 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Elysium: Exploring object-level perception in videos via mllm
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 60cfe8a5-9937-4b5d-bc77-ddbc1041062f · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dc90eb2-36a2-4d1b-b069-d5b2e2e544b9 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adc8c382-9e94-4f56-85ec-ca0908eea582 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Q-instruct: Improving low-level visual abilities for multi-modality foundation models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 73ae1d19-5b5a-43b0-a087-b1761313f326 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM FreeVA: Offline MLLM as Training-Free Video Assistant
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba097623-f394-4d1d-b78a-5d56ba900150 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Next-qa: Next phase of question-answering to explaining temporal actions
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3044a007-2726-4ce9-8d5b-632948f1254f · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Video question answer- ing via gradually refined attention over appearance and mo- tion
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87d0de7b-e66a-447e-bf0d-8fa8f8cb0c7b · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33b63ce7-df92-4764-8b71-74501644df19 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Ad- vancing high-resolution video-language representation with large-scale video transcriptions
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4c4fdf62-eb77-46d6-9424-feda94ec6a98 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2f6514e6-704e-4001-b46a-e8d3ca9252b6 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Sigmoid loss for language image pre-training
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c19181e4-2406-466c-a473-2748807052e6 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 790fb985-21c2-43ed-b0e6-963eca332aad · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd321899-ed2d-4950-a8d7-7c2f4d6367e1 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Long Context Transfer from Language to Vision
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e88e98d9-66a4-404a-858b-53b2bfbb3211 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64d8fbe5-251c-48ab-a7ee-1d47237aa9d8 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b040157e-9a33-4a45-a7e9-9789ff492ae9 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Llava- next: A strong zero-shot video understanding model, 2024
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e7f489b0-07a9-4e3e-b889-79c40955f131 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM MLVU: Benchmarking Multi-task Long Video Understanding
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fcb3e66-8bff-4fcb-9776-a3c1bd75ce33 · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c48134e-fed9-4550-8d3e-278cc7b834ec · outbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Visual7w: Grounded question answering in images
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 559084cc-f064-4cc9-8f56-f56b9d8f608a · inbound
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 172834a1-4ff0-44bd-980d-4158be0247b0 · inbound
Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b4359393-9c7c-47fa-85a1-44bdbd4b6697 · inbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5449c859-83cf-4de4-87c8-6b5172039234 · inbound
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.