Pith. sign in

Paper Citation Record · LEDGER

VoCo-LLaMA: Towards Vision Compression with Large Language Models

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2406.12275.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.12275 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:55:03.218092Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T12:55:40.367899Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 49af9acd-90f2-4ab2-8674-a7dcf84265aa · inbound

Efficient Multi-modal Large Language Models via Visual Token Grouping cites this paper.

Efficient Multi-modal Large Language Models via Visual Token Grouping VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.386323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.386323Z digest=sha256:8e11afe3320a384bbaea337faa024de7adde152c5eb3632bfd7968edc307a553

Observation 809edccb-1c91-4a68-be7a-c59c4b8baf20 · inbound

Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction cites this paper.

Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T05:21:10.954733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:21:10.954733Z digest=sha256:8dfdc4c35d6de9c605817910081fca071341fbaf75af96278c928424822ec0b3

Observation f352bcc1-23e7-4d6a-ade1-4a0ad0759621 · inbound

Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification cites this paper.

Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:00.359007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:00.359007Z digest=sha256:b686ec91996485dc24cfd1ed1e92fafeec7272bae7dee9602dd07171f3d8c1c9

Observation 9307404b-c882-4586-aeef-3fab82364977 · inbound

Ponder & Press: Advancing Visual GUI Agent towards General Computer Control cites this paper.

Ponder & Press: Advancing Visual GUI Agent towards General Computer Control VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:29.638163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:29.638163Z digest=sha256:b1f7bbf8bcda1d077253a7fc8ed688c8db3df8048a3e1b015725f6816313afb1

Observation 94e33742-a456-418b-8ce9-4b18cac4efd5 · inbound

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs cites this paper.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.370458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.370458Z digest=sha256:ffe2e62189c78e54bb40888221c5658e49b13f4fea71211f439ee8cb1c188b4d

Observation 6bcebf37-bfbd-47d8-be5d-00105a08efd5 · inbound

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs cites this paper.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.713337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.713337Z digest=sha256:215fe9defe99109371800e93c4beb5c77cd1ff0919a6638fcacc0da87406b75f

Observation bbb1f662-c167-462f-b560-654daa488e16 · inbound

ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming cites this paper.

ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T23:39:18.195610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:39:18.195610Z digest=sha256:e4e9d91ee81ab2ba4aa00c494eaf4c479d24c34b5aa8a3187d3d4efc20359b3c

Observation 08ff8a72-db80-4fa8-8245-9c42b0332b51 · inbound

DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs cites this paper.

DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T10:55:03.218092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:55:03.218092Z digest=sha256:c1b9c00a32b9b05a3431c048094da9c60f0240a1300f13591b4334a5695eadb3

Observation 0d1d4fab-a1c0-4627-895a-05d81901127d · inbound

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning cites this paper.

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:56.264830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:56.264830Z digest=sha256:99d3137e0f9d7272deb466da61c4d65392462b8a85e0d00b15310cd5eb7dff98

Observation 03a1e32b-abe3-466e-a411-0abe0a4039db · inbound

Towards General Continuous Memory for Vision-Language Models cites this paper.

Towards General Continuous Memory for Vision-Language Models VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:44.748456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:44.748456Z digest=sha256:ee1b9147c57af48c10540bc2542fc9bba692a6be1463bf2b5ec00ffdbf2c7dd4

Observation 949cbecf-4b31-44f1-933a-4ae41c6f8d5f · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:55:40.369819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:631a0876c2807f3adc854887646488b53ff513dfe668e82c85e13816995cdf3f

Observation 085c94df-aa16-4d8f-926b-7d6a5fec40e2 · inbound

CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models cites this paper.

CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:17.253571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:17.253571Z digest=sha256:566dde6a41ff24cd4b7cfa914b848b227b285d6099dfca2bd2e83c6ca8e3192b

Observation fa0c8107-91a4-4810-9bac-fc901fbcce76 · inbound

GEM: Empowering LLM for both Embedding Generation and Language Understanding cites this paper.

GEM: Empowering LLM for both Embedding Generation and Language Understanding VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:50.940626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:50:50.940626Z digest=sha256:c61a8c1c7f86206b13562bceb24228a08ecd194b45a5b68951c06151b0800fe2

Observation 68ed76ec-112b-4f43-92a6-a1c7b91af129 · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.267392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.267392Z digest=sha256:42066c2c86611af201891fd863a1d6f43817d96cc6cfcbbd551e19a19084c6e5

Observation a52528c2-d11a-43dd-a239-7af0352ca0dd · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:35.303741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:35.303741Z digest=sha256:b9a9cc58ec13c8c14524668ee666e58b97592c639109475e43f4111fdedba614

Observation 451ae934-c65c-42c8-9c5d-bad3c4c82881 · inbound

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning cites this paper.

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:39:07.070965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:39:07.070965Z digest=sha256:9ef69dc1845b4671d20639650a1505c9965e96bbf08f0f7ec802d05ffe444f64

Observation 84c13542-3f1b-49c4-b010-e4b1694cf24c · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.529576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.529576Z digest=sha256:45d61eea09cb0b2e1f36d0fd3005230307aeaa7557d83e689afcbee7b70a0b29

Observation 876f6bfa-4589-495d-878e-6a2f0db7cf72 · inbound

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning cites this paper.

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:34.402360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:57:34.402360Z digest=sha256:7567afd487ce3e588e8cd5c732ab8ede2a280e2a201db25b6aad968e131c7ce8

Observation d3e064d2-1d8b-4cfb-8768-b50960fa318b · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 201

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:99fa9a5e683257e70e707afc943ab7cfa5ec4abb72111a208e8c222f352c0a33