Pith. sign in

Paper Citation Record · LEDGER

VoCo-LLaMA: Towards Vision Compression with Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2406.12275.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.12275 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:56.264830Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T12:55:40.367899Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0d1d4fab-a1c0-4627-895a-05d81901127d · inbound

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning cites this paper.

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:56.264830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:56.264830Z digest=sha256:f7c6c170dc3f9332d67665b1db93807ce861ebb6fefd249d9cc7b7801cab06ed

Observation 03a1e32b-abe3-466e-a411-0abe0a4039db · inbound

Towards General Continuous Memory for Vision-Language Models cites this paper.

Towards General Continuous Memory for Vision-Language Models VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:44.748456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:44.748456Z digest=sha256:4dd1f9f53837442ba3082b2848c44d12aeb342cf548182d23e7834d62d54a06f

Observation 949cbecf-4b31-44f1-933a-4ae41c6f8d5f · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:55:40.369819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:193c96efee106d6a1977f5dd264dfd533657f9348d861b12fb377eb427d1bfb9

Observation 085c94df-aa16-4d8f-926b-7d6a5fec40e2 · inbound

CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models cites this paper.

CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:17.253571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:17.253571Z digest=sha256:838447440675417882311c7e13429adaf82ce60f006daedec3cb570940153d36

Observation fa0c8107-91a4-4810-9bac-fc901fbcce76 · inbound

GEM: Empowering LLM for both Embedding Generation and Language Understanding cites this paper.

GEM: Empowering LLM for both Embedding Generation and Language Understanding VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:50.940626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:50:50.940626Z digest=sha256:d73db9bd51d968d57591cc30d3ba9fa86687704cab4c58e95aafa226feeaf2ac

Observation 68ed76ec-112b-4f43-92a6-a1c7b91af129 · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.267392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.267392Z digest=sha256:fae5fa3ff263e1f73e877eabb03a5257ff59a1b0ead9bef9975251ee4649e367

Observation a52528c2-d11a-43dd-a239-7af0352ca0dd · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:35.303741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:35.303741Z digest=sha256:54acbe2188615839e8b69e5030e0b461236740a7ff0e08e629d136a4507d11b9

Observation 451ae934-c65c-42c8-9c5d-bad3c4c82881 · inbound

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning cites this paper.

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:39:07.070965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:39:07.070965Z digest=sha256:33749368cabcccd7008e9d4c8f377b40e8cdd4f7cc096a434c7adc6d7b6d35cd

Observation 84c13542-3f1b-49c4-b010-e4b1694cf24c · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.529576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.529576Z digest=sha256:6419d336ac6e0bb04474059ba761f9a7c1de588599ab68c5a767264f49752025

Observation 876f6bfa-4589-495d-878e-6a2f0db7cf72 · inbound

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning cites this paper.

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:34.402360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:57:34.402360Z digest=sha256:b87ac03eaaae13e7baf973cae4fbb1a71fb5624a1dcb1fb5bea4174fdb508c1b

Observation d3e064d2-1d8b-4cfb-8768-b50960fa318b · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 201

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:83f567f3bfa1174aed2b246d6e85df1e588124eb3c0070e96495ae00a7359d83