Pith. sign in

Paper Citation Record · LEDGER

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models

As of 14 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 3 inbound Pith citation observations for arXiv:2507.02279.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02279 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:38:32.385290Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:51.573559Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T23:04:01.690792Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 45eb87ab-6ff5-493e-a973-cf7ec64bea59 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:27.564674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:27.564674Z digest=sha256:a07e00f23f6e5bb0f9442a443b88438a528be9404baffd6e4199795de419e7b5

Observation 96eb1822-64cc-4d7c-84a2-ac44ece64baf · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:36.847545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:27.608723Z digest=sha256:fa578477033ef7fd16b178af6570b535c0bf72637d0ee2383a1b11a11cad359e

Observation 7d6fbd43-f331-4876-80a2-74e42192b808 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:27.710290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:27.710290Z digest=sha256:230e644a6d377378cb1d7bd672a0d2a63d93998d18b0013c494609ad7a52b241

Observation 56be8d51-f16c-42ba-8ff6-0a3146a356c9 · outbound

This paper cites Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yuxin Pan, Giancarlo Baldan, and Oscar Beijbom.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yuxin Pan, Giancarlo Baldan, and Oscar Beijbom

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:36.660592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:27.786377Z digest=sha256:76d14e3f1ea7c57fd8ecaed634234d3c36ce2887dc0ee54d04055a9c1f701b56

Observation 77b8ae74-2198-4544-ae4a-36cf17682180 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:36.463665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:27.867184Z digest=sha256:3f5222c776a3bc03553682dc6ad4a3cd69f5df3d3d0ccd6d10f8a834943b758c

Observation 4bd692fd-34dc-4c9c-b8cf-9769beced048 · outbound

This paper cites Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:27.955343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:27.955343Z digest=sha256:2e5f29e2d2b67b7db6f8641646214b53b029530a0c90e4e97437a7b3864a9d0d

Observation 6613ff63-f640-43a2-9ede-d4c90e2959af · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:36.279083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:28.050513Z digest=sha256:7c0642cb7c2da2ace7cbba442ed3cb2ed109c10ec42778392fdf5215f3439332

Observation 75d88921-2c27-4bab-85c8-9e2c495c3aa1 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:36.099822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:28.116856Z digest=sha256:bcc5cf8d631a6e299b4ba243d033a8b9c1554be146851e7d1d9c75178570589f

Observation caede71e-f96d-48ed-a41c-f6beb86262e0 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.170493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.170493Z digest=sha256:8454d4a0661496c28e1f596e793b05faf6ae17fe5efb234750a2867415cc377f

Observation e1e7c77a-1365-4953-b547-1bc2f226a3bb · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.234804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.234804Z digest=sha256:5304fd3aeb7d4b60ff2bcd3685e5e35d1d613b22fa95cef633c3bc5453ac37bf

Observation cd48b81d-8482-4fdd-a5c7-a0573699a3de · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.279660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.279660Z digest=sha256:edb618777963177ae4e2ba9a683e1babea7c39530b3276ced0041970806e6ca7

Observation bd4da050-87f3-4fbe-a149-ff457942a408 · outbound

This paper cites PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.365050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.365050Z digest=sha256:7e6f30052058bac88d9f095757a1729c8791fe76c106ed0ecd2d48bf036f6b68

Observation 6f6ae646-e0b0-4d41-84a5-e53dd7dd9545 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.429443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.429443Z digest=sha256:ffc2024f3371deb2f6da0fc1ed112965a6d3c32b4bd0bf22d848438408550f5e

Observation bc6868f7-3f3d-4eb3-beb1-bd800e08d314 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.510632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.510632Z digest=sha256:cb3aff5dbd0b93f90fb4b57c28a70acec74f152b1f78648e6252ba05cd4167d5

Observation c3e5d3dd-996c-4798-870b-8f3b917fb0bc · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.573025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.573025Z digest=sha256:63e144bb71ab1e2744b2c14f8128bbcf32c9452338144d407c242d3a336fbd45

Observation a26c49c9-2976-4085-b180-0cb9d3af72cb · outbound

This paper cites Multimodal Autoregressive Pre-training of Large Vision Encoders.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.632831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.632831Z digest=sha256:d29589a453106b7004e8d554301272357668b46be0692fc68d5916614d5cd6a0

Observation d9b2e771-522f-4535-b467-592bbc13bcc1 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.675201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.675201Z digest=sha256:d1f39e043dcc195b45681c01dcfd1b76612e63bcd967aa2bb1e9ad1670d22c4d

Observation e1a4f98a-1838-4266-8391-2cfa32afdb38 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.730737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.730737Z digest=sha256:5611081eb9584bcb249aab28eed80a8ec397c3596f76192ab2c968fd1fca1ed3

Observation f6c081ea-a824-4099-9d7e-dffb9f08d1b4 · outbound

This paper cites 3D-LLM: Injecting the 3D World into Large Language Models.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models 3D-LLM: Injecting the 3D World into Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.791020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.791020Z digest=sha256:2cf91c5bf1b747d6c680fafe1ea57ed0d6dd21f061d867f42273c0e1d0fa541d

Observation 684a27d0-e59a-4b31-838d-91b7f8f1041e · outbound

This paper cites Learning to Describe Differences Between Pairs of Similar Images.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Learning to Describe Differences Between Pairs of Similar Images

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.839944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.839944Z digest=sha256:a1caa48260bc24ef3620bc6c60a352fb8fecf1348b2cd993f8339f481b28c084

Observation 96fb4c26-88ca-4001-acb5-51d0716555fa · outbound

This paper cites Gemma 3 Technical Report.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Gemma 3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.915132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.915132Z digest=sha256:8f0532fc0a9dc16069ab55aff60eee89a9c926a71b6028cf09af297e947a9f4f

Observation 18b1bc50-ac44-491b-acfc-06a6ae011c74 · outbound

This paper cites A Diagram Is Worth A Dozen Images.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models A Diagram Is Worth A Dozen Images

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.998815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.998815Z digest=sha256:5b509828df312a753d36a7fd5df6c6ee6241a6dc8e79094d2a3fc8c559e46a3c

Observation ea5fccea-bfb3-473f-a75a-2b4aa61e554d · outbound

This paper cites AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.091565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.091565Z digest=sha256:4bb5e4678f6125db36232aa59f52f1e6d3bd82a1d916be8f08992d4d0467d537

Observation a9809df1-5b6a-4630-8349-c8edef12dc28 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.185593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.185593Z digest=sha256:325f9956411742a9a55750feeae0381e1838fd77bbf5795712012a21250019fc

Observation 01d5f8f4-2618-49d7-8e6d-ff9664f1b977 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.262223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.262223Z digest=sha256:3ed8ea5bb9b2616dbd15581abce5036adadf0b95c1df844f2cc47eb5a0aa89a0

Observation f51bc516-a37c-4429-bd80-b08dcedba05f · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.356941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.356941Z digest=sha256:5a422ada7a9ef39bf2813787c3b6f0a5a7ca808fac841aa31f87504137d376bf

Observation 34958276-5601-46d2-a9fe-fc054773bdba · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.433115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.433115Z digest=sha256:06c2e15e7f57d50789593114852d2411de1340b93702b9bd877a0722752c8917

Observation ff1baa90-96ec-421e-b2f2-4d87c4cdd1f2 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:35.974131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:29.493720Z digest=sha256:a16a8ec52fe656a723fcc3a1540088ab6b0f13c4ff5f94b7fe7d06ad6606ae60

Observation c934db64-3e50-42a6-b795-2596f8f744e4 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.543299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.543299Z digest=sha256:d909d1bfdd07a2b5cde3e987258163913049e43e5f4884d090114d69c5d0a46d

Observation 3402898e-e1b4-46ec-bee6-7ef9ab791b52 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:35.857868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:29.601889Z digest=sha256:07f4c86b2328543405198c10183cc9bc3c6f3a13bfb4685226d0cd32adfc20b9

Observation f2a2ffec-66df-4cda-a57d-fab3e0840b24 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:35.670119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:29.707324Z digest=sha256:fedef59535ab812d3f2ba08fbb8c1da397c6fcbbef015fa7e35902778ac2b136

Observation e87ff870-e169-4340-b05f-d5365e1ee444 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.802220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.802220Z digest=sha256:d479e6c2497537cc48f60a628db65f072ea4a405e62e619bdd2d0f7845d8033a

Observation 9ce3b2bb-fc0d-4ffc-9dd8-2b404bf4b475 · outbound

This paper cites Visual Instruction Tuning.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Visual Instruction Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.866140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.866140Z digest=sha256:ba3819a21ab280f9a704304ef5716e3930f3a5ab06cd5e8060018e6efd2eab6c

Observation 6e178c9b-eb14-4504-aba2-7da9adfe3b1b · outbound

This paper cites What Large Language Models Bring to Text-rich VQA?.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models What Large Language Models Bring to Text-rich VQA?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.985659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.985659Z digest=sha256:942ab4d4f78572809f6345a55a07b483100c08e54e829b82e54cf0d105ff8470

Observation 8b7a0d1f-f32b-45fc-bde1-3cc079cb8293 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:35.424039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:30.023977Z digest=sha256:389bc689c66f46891866f45cd697e87788c6e8e394b2ef13a21edfed36dc2d05

Observation 48ae9154-6455-4740-9750-a32c2ad19818 · outbound

This paper cites InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.064975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.064975Z digest=sha256:bba839b951e5c833dc9d3a7ebd20fdbdc1adb0d7b14de2ab6bdd48d874aca52f

Observation 468d22ad-8ab8-4b9f-b59b-9d312cb0db29 · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.134630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.134630Z digest=sha256:66b55308bbcca3bb9d961c22bc3f4598be93d0f9744217ad1321880bdcbaf87b

Observation b1200b95-bdb6-409e-ab62-b0f44ef5bc3e · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.194895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.194895Z digest=sha256:4cdb86422c0198dc6c6afd276ba21c29fb40e1ac763319fdc7e6e7255abd0477

Observation 23e8eb76-4365-4923-9169-307a01d86f3e · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:35.163832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:30.266880Z digest=sha256:719ba99b7414d05d6f0a1b04a51fea6030deb2c17446151b1f7e7e8fcab68c40

Observation 430a8010-0505-408a-a53b-b1ddfc928d89 · outbound

This paper cites Manmatha, and C.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Manmatha, and C

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:34.905992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:30.391497Z digest=sha256:45d6d282ba31374740ef42068784eb4423e3f4dbd36953a3914a14705c0b0a66

Observation bdd09533-6b82-46f0-921e-33d77498e756 · outbound

This paper cites Perception Test: A Diagnostic Benchmark for Multimodal Video Models.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.481104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.481104Z digest=sha256:cb64fbac410aa8d1b158c003ebb22f929c67513ebf24e8de22415f43ac4ae33b

Observation f0c5378c-1977-4616-9ceb-5d27526d1c01 · outbound

This paper cites Multi-Image Visual Question Answering.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Multi-Image Visual Question Answering

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.602789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.602789Z digest=sha256:3ce238ef28179b3453c6f24132fce781e589781533e0280abc83430d43bc7d8c

Observation f07f6cfb-786e-498d-83da-39dd7c4785f2 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:34.718389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:30.689621Z digest=sha256:d3cb086c2a811deeaa5bcfa1788de785f0ecc605b2b9a794177f2226c6ea8b32

Observation 067556ee-575c-437a-be1d-9ef6674c6a55 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.765419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.765419Z digest=sha256:dedc4c1e648dd36f793d2c1841b93b4eb382a4daf08b55a687102c99b1bc9423

Observation 7e27a2d7-5a94-40e4-92bb-6e9d5a6a233a · outbound

This paper cites Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:38:34.482370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:30.808494Z digest=sha256:497f259cec1b19427313a1e52f1e180f90be0169666c5a7d1cd59b881cd9ac09

Observation d92d97ae-7ffb-4016-9bdb-12e8a7830f2e · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:34.257481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:30.870943Z digest=sha256:51acd39d4577f03de3bc319181f0bfe4e31ed10ffaf1c1fbc5f916d640a6f998

Observation 7482c423-93bf-4643-bbc4-1e277a0f76ca · outbound

This paper cites NLVR2 Visual Bias Analysis.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models NLVR2 Visual Bias Analysis

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.958476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.958476Z digest=sha256:223c8ad8fb260aeb5a205f118d4dd12abb055cfb33b875142620eb5d5fa6a3a4

Observation b850ff87-f508-44f5-92d4-3fc5b3d01f4d · outbound

This paper cites Visual Storytelling.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Visual Storytelling

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T20:38:32.685743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:31.090662Z digest=sha256:daf778f1bc0a7cf300a98cd2d736da2c09fe627325d8ae43f6009ea1d822daa2

Observation 7c35b610-cf1d-493e-9a50-fdb95330df3d · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.214733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.214733Z digest=sha256:a3d002ec8d3f7af95333715b18db76a4997abc7b0416f4eee643c50e8051e05d

Observation e360d848-1f3e-45bf-bdc2-152f0015cf5e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.337702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.337702Z digest=sha256:bbc08ec19011333f9c584fb15db4beca361607339562f532973f8310c224d8e6

Observation a47cc575-87af-4819-83cc-515df417841f · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.429497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.429497Z digest=sha256:22ef3cb9e84ec1022c6b50dbfc0f203a56eacc295a6fabfdebededa28872fe0a

Observation a3c78e87-7777-4fb2-900b-a218944e0aef · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.498256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.498256Z digest=sha256:a2624cab260b6b01cffea0e5a72e5e854901b965d6c03d0a22dc4e7bc8f7267d

Observation 2f4c5067-3c3c-4df1-a3a6-4bfeb82af34f · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:34.028934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:31.606539Z digest=sha256:bbd8f197d579792037e60449ab434315c27bd67a2852d3b1eb98ecd0f78d07ea

Observation add12b7a-b9fe-45c0-9bb7-5b3f252d9128 · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:33.776607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:31.676473Z digest=sha256:519e0277b77856c6065c4d78a84cf1278fcf7a59d0a1954fa0bfd7195ab61a7d

Observation 35d14d73-0835-4788-9491-3e49152ba743 · outbound

This paper cites Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.752838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.752838Z digest=sha256:33c3cefbd93b4e428e2ad6fc9bba7c31ac0a270aeaa0aa4fb730db5a4d486cc3

Observation dc5ccaf7-9a03-4a84-91d1-53fc6ca88fb7 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.911548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.911548Z digest=sha256:34d4b989bb935bbd1d6bf4eca612ae15a38f99edf3e3afcc2b9ce5f4330877ab

Observation b038f03d-4685-4341-9e43-c500f903b73c · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:33.582663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:32.007214Z digest=sha256:b8889179713c253cbed2a5064885657c01552d607b74517c3fc05e8e99590a87

Observation 2329d619-41dc-479a-beb6-cceeeb518c2e · outbound

This paper cites an unresolved cited work.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:38:33.337407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:38:32.115722Z digest=sha256:d3a30b3c5a53b355b7366c1a85934959ee28d5e829d9fc3b2072f2856c352c4c

Observation 2041be15-4664-41e7-ab12-8b3cc899a4d5 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.171118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.171118Z digest=sha256:5a90a3346ef641f49a1802e60c89274e7fd5ed623bd1bb24935598e14e8caffa

Observation 7fb6840f-d5aa-4419-bcfa-772a6a478385 · outbound

This paper cites online" 'onlinestring :=.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models online" 'onlinestring :=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.267160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.267160Z digest=sha256:8791d9830357b83de9baf56306ac5a3bae13e1ea5cf0591e4518fec534e175e4

Observation 0d020b1b-b650-46dd-942d-733f1c32d85c · outbound

This paper cites write newline.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.385290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.385290Z digest=sha256:1cad0b703b98c40c4da91dba5bd3d5adcc035e0d73e200000d3a7216b78338a1

Pith citing papers

Observation de24ad38-9c42-4984-a9d0-dc8f4eec8aa1 · inbound

Stateful Token Reduction for Long-Video Hybrid VLMs cites this paper.

Stateful Token Reduction for Long-Video Hybrid VLMs LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:05.183330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:05.183330Z digest=sha256:049a518dc60d17eb8f93fd9e2a5ade00ef66580ba994918fa6c5370ae170dd7f

Observation 181c232c-4435-47c8-bb9d-5b6f28653334 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models

Reference 207

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.692379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:f15c9dfb21d155dddffe228bf98b8ace45e49f03ff1274ade8db635667896cc7

Observation a3c4e44a-af0c-4238-9ed0-9d3ae8b95f91 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.573559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.573559Z digest=sha256:1764cf083f8a40724fb35eee9bb635acf51465145503e5a4632547076fc5e425