Pith. sign in

Paper Citation Record · LEDGER

Efficient Large Multi-modal Models via Visual Context Compression

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2406.20092.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.20092 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:11:32.736571Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:48:03.161959Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 46f9e84c-7eac-4789-9d20-70dc9040a77d · inbound

PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction cites this paper.

PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction Efficient Large Multi-modal Models via Visual Context Compression

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:12:14.743952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T12:12:14.613620Z digest=sha256:576ad9a9e0011cc5fca3a34a68301ca575fe7a12dacc93a5ba49f46a3f3a8a86

Observation dd31a8fc-f16e-4dcd-bc52-ab91dbd3186a · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Efficient Large Multi-modal Models via Visual Context Compression

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.682863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:a67f09db344e1efd6f0a70e0d1bfa8e527b9eec728c81884f1b2687a45a41e0c

Observation 5c4ded37-74fb-480f-a64a-039396ba00c7 · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models Efficient Large Multi-modal Models via Visual Context Compression

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:02:17.873890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:65d8aa5d27e406405e7b53d001fbf9811517c600062469f2d849b0ddb9b94a72

Observation 2167d069-857a-41eb-9072-65fe975ff2e5 · inbound

Efficient Multi-modal Long Context Learning for Training-free Adaptation cites this paper.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Efficient Large Multi-modal Models via Visual Context Compression

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:32.736571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:32.736571Z digest=sha256:768c45015e2272d4e634ba423d7bdde735d35b1621d647ede39e2839b318cae2

Observation 28c77754-5715-4cfa-a80a-afb162c9921b · inbound

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models cites this paper.

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models Efficient Large Multi-modal Models via Visual Context Compression

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:53.603924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:53.603924Z digest=sha256:e8052a0339d7822bf48113ed678bb4c81b9731b1d68ead2260d0fdeafc1ef011

Observation 0e65442a-8ef8-44e2-86c1-fa56cf9150ee · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Efficient Large Multi-modal Models via Visual Context Compression

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.561451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.561451Z digest=sha256:9647810110a0598fcc116595096b4e067ed288803a0f95c9611cdd8b5170e5f2

Observation 9366842e-7f78-4ba8-a8a1-f0bdb62c1076 · inbound

Structured Prompting and Multi-Agent Knowledge Distillation for Traffic Video Interpretation and Risk Inference cites this paper.

Structured Prompting and Multi-Agent Knowledge Distillation for Traffic Video Interpretation and Risk Inference Efficient Large Multi-modal Models via Visual Context Compression

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:39.208373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T19:06:39.208373Z digest=sha256:19a96b70761626260e1c9e10770fbc65a21cc2b4db7aa52027b2cf401d76bffa

Observation 4ca98475-15be-4e2d-b40d-3420c76b12b1 · inbound

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding cites this paper.

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding Efficient Large Multi-modal Models via Visual Context Compression

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:31:01.569823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:50:37.022338Z digest=sha256:59c78cfc999dbe1d6e2881d951e53a044dad89aee405b5f6ab11450cfbbe65dd

Observation 97a43270-b6fe-47e1-b35c-e6153636e3bc · inbound

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning cites this paper.

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning Efficient Large Multi-modal Models via Visual Context Compression

Reference 210

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:07:48.419690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T10:28:11.440915Z digest=sha256:80a812bac0712c579ee51d3fff599ab7717a477b040fee373c458c715b4ea90e

Observation b00b2214-67db-454d-bb4b-8352925e2055 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Efficient Large Multi-modal Models via Visual Context Compression

Reference 295

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.163310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:533a1701ac83615e361a714961c86e4be6ed53861581fe8b3c1fb01105dae366