Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:38:32.385290Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 3 inbound Pith citation observations for arXiv:2507.02279.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:38:32.385290Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:51.573559Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T23:04:01.690792Z
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 45eb87ab-6ff5-493e-a973-cf7ec64bea59 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Flamingo: a Visual Language Model for Few-Shot Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96eb1822-64cc-4d7c-84a2-ac44ece64baf · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7d6fbd43-f331-4876-80a2-74e42192b808 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56be8d51-f16c-42ba-8ff6-0a3146a356c9 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yuxin Pan, Giancarlo Baldan, and Oscar Beijbom
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 77b8ae74-2198-4544-ae4a-36cf17682180 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4bd692fd-34dc-4c9c-b8cf-9769beced048 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6613ff63-f640-43a2-9ede-d4c90e2959af · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 75d88921-2c27-4bab-85c8-9e2c495c3aa1 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation caede71e-f96d-48ed-a41c-f6beb86262e0 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1e7c77a-1365-4953-b547-1bc2f226a3bb · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd48b81d-8482-4fdd-a5c7-a0573699a3de · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd4da050-87f3-4fbe-a149-ff457942a408 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f6ae646-e0b0-4d41-84a5-e53dd7dd9545 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc6868f7-3f3d-4eb3-beb1-bd800e08d314 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e5d3dd-996c-4798-870b-8f3b917fb0bc · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a26c49c9-2976-4085-b180-0cb9d3af72cb · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9b2e771-522f-4535-b467-592bbc13bcc1 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1a4f98a-1838-4266-8391-2cfa32afdb38 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models BLINK: Multimodal Large Language Models Can See but Not Perceive
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6c081ea-a824-4099-9d7e-dffb9f08d1b4 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models 3D-LLM: Injecting the 3D World into Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 684a27d0-e59a-4b31-838d-91b7f8f1041e · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Learning to Describe Differences Between Pairs of Similar Images
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96fb4c26-88ca-4001-acb5-51d0716555fa · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Gemma 3 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18b1bc50-ac44-491b-acfc-06a6ae011c74 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models A Diagram Is Worth A Dozen Images
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea5fccea-bfb3-473f-a75a-2b4aa61e554d · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9809df1-5b6a-4630-8349-c8edef12dc28 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01d5f8f4-2618-49d7-8e6d-ff9664f1b977 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f51bc516-a37c-4429-bd80-b08dcedba05f · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34958276-5601-46d2-a9fe-fc054773bdba · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff1baa90-96ec-421e-b2f2-4d87c4cdd1f2 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c934db64-3e50-42a6-b795-2596f8f744e4 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3402898e-e1b4-46ec-bee6-7ef9ab791b52 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f2a2ffec-66df-4cda-a57d-fab3e0840b24 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e87ff870-e169-4340-b05f-d5365e1ee444 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ce3b2bb-fc0d-4ffc-9dd8-2b404bf4b475 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Visual Instruction Tuning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e178c9b-eb14-4504-aba2-7da9adfe3b1b · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models What Large Language Models Bring to Text-rich VQA?
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b7a0d1f-f32b-45fc-bde1-3cc079cb8293 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 48ae9154-6455-4740-9750-a32c2ad19818 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 468d22ad-8ab8-4b9f-b59b-9d312cb0db29 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1200b95-bdb6-409e-ab62-b0f44ef5bc3e · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23e8eb76-4365-4923-9169-307a01d86f3e · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 430a8010-0505-408a-a53b-b1ddfc928d89 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Manmatha, and C
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bdd09533-6b82-46f0-921e-33d77498e756 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Perception Test: A Diagnostic Benchmark for Multimodal Video Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0c5378c-1977-4616-9ceb-5d27526d1c01 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Multi-Image Visual Question Answering
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f07f6cfb-786e-498d-83da-39dd7c4785f2 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 067556ee-575c-437a-be1d-9ef6674c6a55 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e27a2d7-5a94-40e4-92bb-6e9d5a6a233a · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d92d97ae-7ffb-4016-9bdb-12e8a7830f2e · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7482c423-93bf-4643-bbc4-1e277a0f76ca · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models NLVR2 Visual Bias Analysis
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b850ff87-f508-44f5-92d4-3fc5b3d01f4d · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Visual Storytelling
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7c35b610-cf1d-493e-9a50-fdb95330df3d · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e360d848-1f3e-45bf-bdc2-152f0015cf5e · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a47cc575-87af-4819-83cc-515df417841f · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models CogVLM: Visual Expert for Pretrained Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3c78e87-7777-4fb2-900b-a218944e0aef · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f4c5067-3c3c-4df1-a3a6-4bfeb82af34f · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation add12b7a-b9fe-45c0-9bb7-5b3f252d9128 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 35d14d73-0835-4788-9491-3e49152ba743 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc5ccaf7-9a03-4a84-91d1-53fc6ca88fb7 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b038f03d-4685-4341-9e43-c500f903b73c · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2329d619-41dc-479a-beb6-cceeeb518c2e · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2041be15-4664-41e7-ab12-8b3cc899a4d5 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fb6840f-d5aa-4419-bcfa-772a6a478385 · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models online" 'onlinestring :=
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d020b1b-b650-46dd-942d-733f1c32d85c · outbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models write newline
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de24ad38-9c42-4984-a9d0-dc8f4eec8aa1 · inbound
Stateful Token Reduction for Long-Video Hybrid VLMs LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 181c232c-4435-47c8-bb9d-5b6f28653334 · inbound
Toward Native Multimodal Modeling: A Roadmap LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models
Reference 207
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a3c4e44a-af0c-4238-9ed0-9d3ae8b95f91 · inbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.