Pith. sign in

Paper Citation Record · LEDGER

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models

As of 9 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 3 inbound Pith citation observations for arXiv:2507.02279.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02279 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:38:32.385290Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:51.573559Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T23:04:01.690792Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 45eb87ab-6ff5-493e-a973-cf7ec64bea59 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:27.564674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:27.564674Z digest=sha256:d02f84a72d594edd4217f336f663718fafddadf066a5b1deabb91414bb81e945

Observation 96eb1822-64cc-4d7c-84a2-ac44ece64baf · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:36.847545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:27.608723Z digest=sha256:f2814adc9de0cdaec05f9fdf7af92bcd9291670c01654a1d8c79fae2697cda89

Observation 7d6fbd43-f331-4876-80a2-74e42192b808 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:27.710290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:27.710290Z digest=sha256:35832ef77ed3c30e5e7f5021049d41f70260b25f5aae48c882262133bbc895ef

Observation 56be8d51-f16c-42ba-8ff6-0a3146a356c9 · outbound

This paper cites Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yuxin Pan, Giancarlo Baldan, and Oscar Beijbom.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yuxin Pan, Giancarlo Baldan, and Oscar Beijbom

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:36.660592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:27.786377Z digest=sha256:2572a3cf87c1ed2c0e90a36afdf8c8fd7df364c579547930a59cfa08e2523d03

Observation 77b8ae74-2198-4544-ae4a-36cf17682180 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:36.463665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:27.867184Z digest=sha256:d31d6f9976e1d2e89788de1c2a0e290dae19365a4ccde7549b0a5ff686aa7454

Observation 4bd692fd-34dc-4c9c-b8cf-9769beced048 · outbound

This paper cites Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:27.955343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:27.955343Z digest=sha256:40603b9c310f0364c928a62ddc059f9905255caac2573e49163fe575e7972b9a

Observation 6613ff63-f640-43a2-9ede-d4c90e2959af · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:36.279083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:28.050513Z digest=sha256:95c9d9a71553ce82c4ed8db939eb6e05ad97747c6f5975714aa9d3c4bb36e059

Observation 75d88921-2c27-4bab-85c8-9e2c495c3aa1 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:36.099822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:28.116856Z digest=sha256:bc78b56f16f10180449b0d734bcec27c71ac4cdd0f00b119eef0c6c6f25039b9

Observation caede71e-f96d-48ed-a41c-f6beb86262e0 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.170493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.170493Z digest=sha256:7f8699847b1b7c822adaaa3a8ef01d8b967fd0d8baab588d938396ceae265534

Observation e1e7c77a-1365-4953-b547-1bc2f226a3bb · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.234804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.234804Z digest=sha256:0380e66fc76c4558cbf607cba0146cc0d704069b5e4879d778f69f20520d7fed

Observation cd48b81d-8482-4fdd-a5c7-a0573699a3de · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.279660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.279660Z digest=sha256:5ae6cd10495f5757c02f4d4ef3e93bfd03e6716efef8cc2befdb5c8943d1fa42

Observation bd4da050-87f3-4fbe-a149-ff457942a408 · outbound

This paper cites PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.365050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.365050Z digest=sha256:42d58c9412d942edab5745c0df734f8aa15c95367b197878899e1c192cda444b

Observation 6f6ae646-e0b0-4d41-84a5-e53dd7dd9545 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.429443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.429443Z digest=sha256:6c395c4de11a823802b45890ca7ec4bcbf279aebf980d5267ea565cd454af3fb

Observation bc6868f7-3f3d-4eb3-beb1-bd800e08d314 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.510632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.510632Z digest=sha256:2a426a03d23a382a7dc4bf285cb008f2556bd660e9fcac4432568d6622500e62

Observation c3e5d3dd-996c-4798-870b-8f3b917fb0bc · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.573025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.573025Z digest=sha256:dfec2938eb81630bdeaade20c9e8eaa071df8a506a3f75ad2caf506932de02f5

Observation a26c49c9-2976-4085-b180-0cb9d3af72cb · outbound

This paper cites Multimodal Autoregressive Pre-training of Large Vision Encoders.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.632831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.632831Z digest=sha256:d1083a1558258afda5878ad317f31aba0a93ef2ea006e27931492c867258a349

Observation d9b2e771-522f-4535-b467-592bbc13bcc1 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.675201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.675201Z digest=sha256:969368e58f77e6bb890b9d1f0df4ee0718b9b04da08ace8830463da9efa83cda

Observation e1a4f98a-1838-4266-8391-2cfa32afdb38 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.730737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.730737Z digest=sha256:dc1679fb130c1a0452eafe3d1088ad9d5ed48aea0e6620a636fde02e8570fecf

Observation f6c081ea-a824-4099-9d7e-dffb9f08d1b4 · outbound

This paper cites 3D-LLM: Injecting the 3D World into Large Language Models.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models 3D-LLM: Injecting the 3D World into Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.791020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.791020Z digest=sha256:0662c1565b622cd1156a71ddd9ef1d763888dd84a6c8f833d439d3a310dd2f0a

Observation 684a27d0-e59a-4b31-838d-91b7f8f1041e · outbound

This paper cites Learning to Describe Differences Between Pairs of Similar Images.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Learning to Describe Differences Between Pairs of Similar Images

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.839944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.839944Z digest=sha256:fab3a2e7b44269775da5c31a70071c10c4f3b76ff1d49ca22e50bd0be74d95f8

Observation 96fb4c26-88ca-4001-acb5-51d0716555fa · outbound

This paper cites Gemma 3 Technical Report.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Gemma 3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.915132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.915132Z digest=sha256:fbb462fffbee6f5f69e005bf84009bf90a0be31d51294e43607973b4681f4bde

Observation 18b1bc50-ac44-491b-acfc-06a6ae011c74 · outbound

This paper cites A Diagram Is Worth A Dozen Images.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models A Diagram Is Worth A Dozen Images

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.998815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.998815Z digest=sha256:670b85026f8fafc29869a0f23818f330c268fb55093b4c8089cc31b39d03236e

Observation ea5fccea-bfb3-473f-a75a-2b4aa61e554d · outbound

This paper cites AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.091565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.091565Z digest=sha256:7c09006c05684405a19ba734fe95488463d9fa2fd829f6b7f5b6978aa4e1710b

Observation a9809df1-5b6a-4630-8349-c8edef12dc28 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.185593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.185593Z digest=sha256:73e707bd32661993aedc05942e6c7a6347dd4defbfd73e6c489ff84928006617

Observation 01d5f8f4-2618-49d7-8e6d-ff9664f1b977 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.262223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.262223Z digest=sha256:aaede87dbd142ddcf2d1eba437881dd91b9261a2b0d071e5ebdfb2b793eb6b16

Observation f51bc516-a37c-4429-bd80-b08dcedba05f · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.356941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.356941Z digest=sha256:9e0cd9bf6e1e047234324669dcfea1c755f8c252e43eb67c920091d93f05133b

Observation 34958276-5601-46d2-a9fe-fc054773bdba · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.433115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.433115Z digest=sha256:89b8e8369fbdb52148e5c58181ff1ac9a4efef2d4bb70c2bd8d280ed8abc86af

Observation ff1baa90-96ec-421e-b2f2-4d87c4cdd1f2 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:35.974131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:29.493720Z digest=sha256:53d730af8b7ea3692fab3347feefac53c30399d94109dce9485f4353816493ff

Observation c934db64-3e50-42a6-b795-2596f8f744e4 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.543299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.543299Z digest=sha256:88a72439f0d16d3fac71826c0bb4218e9d1f0acbe32ce5588311a5616ce68bba

Observation 3402898e-e1b4-46ec-bee6-7ef9ab791b52 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:35.857868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:29.601889Z digest=sha256:249e57cde2272eac450f4c58eb5bee105f04ac3d6c51c6737e5e5cb0faec58ac

Observation f2a2ffec-66df-4cda-a57d-fab3e0840b24 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:35.670119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:29.707324Z digest=sha256:ad24add2a3f1c3a7647a1af7fefd94a9588b52cbe89a0a9d05a127e0fba30f50

Observation e87ff870-e169-4340-b05f-d5365e1ee444 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.802220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.802220Z digest=sha256:f26f3a31b266a1cfb1d23ad7a859a554d6ae5e4c5e4c72bf44689c894cb3bbc8

Observation 9ce3b2bb-fc0d-4ffc-9dd8-2b404bf4b475 · outbound

This paper cites Visual Instruction Tuning.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Visual Instruction Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.866140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.866140Z digest=sha256:5b3c728007134f1e420f43ac08599513ea57b30ad33f3483aa9bc046f5902bd3

Observation 6e178c9b-eb14-4504-aba2-7da9adfe3b1b · outbound

This paper cites What Large Language Models Bring to Text-rich VQA?.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models What Large Language Models Bring to Text-rich VQA?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.985659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.985659Z digest=sha256:3b37d95bedd6d18a8b821f742107ef8aa7a9c94e02ac98dbd5996c0f90f1d427

Observation 8b7a0d1f-f32b-45fc-bde1-3cc079cb8293 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:35.424039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:30.023977Z digest=sha256:262a84bc21f0e47f1cb8e0035e19dcdade2aaff366f00dd7c0e9cd6596dcebc2

Observation 48ae9154-6455-4740-9750-a32c2ad19818 · outbound

This paper cites InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.064975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.064975Z digest=sha256:c18481eb5525c303802d7bc7cc614dfa0f0ee3f720bf7590ec1cb535d1b045c7

Observation 468d22ad-8ab8-4b9f-b59b-9d312cb0db29 · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.134630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.134630Z digest=sha256:262b25f7a64dfc6a389ddf5a3442de84d2f57ff92d855cd66918ddf23a287967

Observation b1200b95-bdb6-409e-ab62-b0f44ef5bc3e · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.194895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.194895Z digest=sha256:8b63ac2a96c46bfa5aabfe78bbb10505d968453a9f83b9cf172db2da25de0aa8

Observation 23e8eb76-4365-4923-9169-307a01d86f3e · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:35.163832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:30.266880Z digest=sha256:d9b2ade2b6de9c4e9c83ce5d0c3f4a4a5add66c8ab271c701ba74056085d0517

Observation 430a8010-0505-408a-a53b-b1ddfc928d89 · outbound

This paper cites Manmatha, and C.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Manmatha, and C

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:34.905992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:30.391497Z digest=sha256:decdaf46e7bf2b0a7577bd3e8f6a517314ec436d44512403afdb0e7ca4cf10c2

Observation bdd09533-6b82-46f0-921e-33d77498e756 · outbound

This paper cites Perception Test: A Diagnostic Benchmark for Multimodal Video Models.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.481104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.481104Z digest=sha256:9b9affabee2e3f99128468ecaaa4801a7b34ed514cd37440a276044fa983e281

Observation f0c5378c-1977-4616-9ceb-5d27526d1c01 · outbound

This paper cites Multi-Image Visual Question Answering.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Multi-Image Visual Question Answering

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.602789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.602789Z digest=sha256:b1d933488ee20cd8d3c874cf62c9b33d70a6df80a97d04308fb1aa4efa627b3e

Observation f07f6cfb-786e-498d-83da-39dd7c4785f2 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:34.718389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:30.689621Z digest=sha256:d3f67b7890d1bc53b8fc417da7b367bf6b66d9878a25c8b2b05de8d8d5efc717

Observation 067556ee-575c-437a-be1d-9ef6674c6a55 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.765419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.765419Z digest=sha256:4d6deb9a40025eb64c041ed752257d8663be6e11e5dfeff0b4b41ae6369b9b77

Observation 7e27a2d7-5a94-40e4-92bb-6e9d5a6a233a · outbound

This paper cites Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:34.482370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:30.808494Z digest=sha256:e690994466c062c9e589d2081f08b3f3c178948367a6f01924d1755ac9a68e14

Observation d92d97ae-7ffb-4016-9bdb-12e8a7830f2e · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:34.257481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:30.870943Z digest=sha256:db72acb0f8455f3a1463b072729c0b9c60dff2cf7ded2a012cfd4f6738624ac1

Observation 7482c423-93bf-4643-bbc4-1e277a0f76ca · outbound

This paper cites NLVR2 Visual Bias Analysis.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models NLVR2 Visual Bias Analysis

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.958476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.958476Z digest=sha256:525b68cc31a1c8072c01b50c88a59113c06454c3521fd0999b75aa07dd24a4fa

Observation b850ff87-f508-44f5-92d4-3fc5b3d01f4d · outbound

This paper cites Visual Storytelling.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Visual Storytelling

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T20:38:32.685743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:31.090662Z digest=sha256:bf94e7d0b4f7a275de18c4654325c7feec34688b703070e501496104e9000951

Observation 7c35b610-cf1d-493e-9a50-fdb95330df3d · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.214733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.214733Z digest=sha256:fbed6800af925a06cdb8f3836c4424e9974320292192677328cb3d83f1ae63b7

Observation e360d848-1f3e-45bf-bdc2-152f0015cf5e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.337702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.337702Z digest=sha256:4b41dc7c34c89eff68758502c75b497f92033ff1203d625e790a173b84d1e8b1

Observation a47cc575-87af-4819-83cc-515df417841f · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.429497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.429497Z digest=sha256:570cd8dc695382f11c9eef86418e4e449235fc29781b052b8283affeac9f61fd

Observation a3c78e87-7777-4fb2-900b-a218944e0aef · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.498256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.498256Z digest=sha256:f75e007c96acc59925cae2409024240f1906e10e4040cb7f6f7ed1e0739411d1

Observation 2f4c5067-3c3c-4df1-a3a6-4bfeb82af34f · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:34.028934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:31.606539Z digest=sha256:a37a7432a1dc08cb584612f9d5b21d4cee42b67173f4ec05d2f6d5a95c060698

Observation add12b7a-b9fe-45c0-9bb7-5b3f252d9128 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:33.776607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:31.676473Z digest=sha256:cccaf24dbbc0f131220ccbef8b2ff71579e1a70d57015be10b64c3d65445bfab

Observation 35d14d73-0835-4788-9491-3e49152ba743 · outbound

This paper cites Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.752838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.752838Z digest=sha256:be7d95b95aa4d4b1f8398aa690e3e9cb5f07b0c39aabc3a7c8cfc8f6ebacdb80

Observation dc5ccaf7-9a03-4a84-91d1-53fc6ca88fb7 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.911548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.911548Z digest=sha256:5857f75f73d26139f6632f20ef44cdb121f25cfb494d6092f624e1415fbc4d19

Observation b038f03d-4685-4341-9e43-c500f903b73c · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:33.582663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:32.007214Z digest=sha256:a0973e3d7e2601baec23e9ee176ab6e119aa25165aaed1ccf44ef11fdba136f0

Observation 2329d619-41dc-479a-beb6-cceeeb518c2e · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:33.337407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:38:32.115722Z digest=sha256:d0ca4f288ae170fe179a6f797ab375b660aeb79c5812ef5275690ba39c4f6955

Observation 2041be15-4664-41e7-ab12-8b3cc899a4d5 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.171118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.171118Z digest=sha256:e1b4a4f774edbafc5eb2f7e444fae86869f2a8743092b59ab915c0c60a8c3565

Observation 7fb6840f-d5aa-4419-bcfa-772a6a478385 · outbound

This paper cites online" 'onlinestring :=.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models online" 'onlinestring :=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.267160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.267160Z digest=sha256:aee1fd1b62b69c2c062eb8cc3551288c19d5df06c94922d7a66d14bc1b14d58f

Observation 0d020b1b-b650-46dd-942d-733f1c32d85c · outbound

This paper cites write newline.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.385290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.385290Z digest=sha256:c6f610dc0230278bbd0bd31fbf67e8630c472ff2a3a5027a0e259485b6c68776

Pith citing papers

Observation de24ad38-9c42-4984-a9d0-dc8f4eec8aa1 · inbound

Stateful Token Reduction for Long-Video Hybrid VLMs cites this paper.

Stateful Token Reduction for Long-Video Hybrid VLMs LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:05.183330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:05.183330Z digest=sha256:5495c8a8de9cc5ab7f2fd3077bc1d6b074443c35fc45fb5b01623d6463a18fc4

Observation 181c232c-4435-47c8-bb9d-5b6f28653334 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models

Reference 207

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.692379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:3f32d5bd0227b3290d5b29dcaf940d4aaa8a128aaa550f24a37f533ab7a41707

Observation a3c4e44a-af0c-4238-9ed0-9d3ae8b95f91 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.573559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.573559Z digest=sha256:5300c094490614183b0c0c0335bb6a793b65c5369d6306f6c12d87fe014d01ca