Pith. sign in

Paper Citation Record · LEDGER

TokenPacker: Efficient Visual Projector for Multimodal LLM

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2407.02392.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.02392 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:47.521580Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:15:44.617819Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5ecbc9d4-5d8e-43c3-a3a6-0753f84636a3 · inbound

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step cites this paper.

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:35:25.953654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T11:35:25.894465Z digest=sha256:ffd4c041f1ef9f77f03f2972548207998d8a386355bd5bf13a9934709a5d6f0d

Observation 2ad7117d-d750-496d-bf95-cb4432bb11c7 · inbound

PixelThink: Towards Efficient Chain-of-Pixel Reasoning cites this paper.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.521580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.521580Z digest=sha256:bc4bdaa06f9b04a79b01144a6220b438db33bc13e164af3cd6a09eef0a6b362d

Observation 9f5a24b9-6c20-4983-ad8f-d9e2bcdde671 · inbound

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models cites this paper.

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:04.930789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:09:04.930789Z digest=sha256:2c5c26f1b4edd428b0c826df4ee820cb0e20449e895da23908d7722231249c1c

Observation 81818563-8738-439b-a096-f401ce88c5b9 · inbound

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding cites this paper.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.333184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.333184Z digest=sha256:0a5d684ceb88c8067c2bb13056b492d27fe4de5c2b3f4285983c37428c8fe716

Observation f1413e60-66b2-4474-8cc5-4e90d4930ee2 · inbound

Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings cites this paper.

Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:48.866029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:48.866029Z digest=sha256:46218405a5fd73ac2ae8da464708acf69c9e0fc985dd827ccb8dfc0e800ee5ee

Observation e8315937-1c4c-4eed-a4d5-020faa7412bf · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:09.538863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:09.538863Z digest=sha256:6d7fc37ddf79bd3ac865d3c49bfd2e68f918cfad4096d1c36bf7804cf12d075d

Observation 4adaa931-5207-495e-bbd1-7f0ca55cb3db · inbound

HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding cites this paper.

HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:47:07.697677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:47:07.697677Z digest=sha256:36a0776f71187efcd95872270ed6fbf82e867670462b4dcc9ff50c3b057de29a

Observation c62664f3-9f31-4c76-a0de-2a28e1bfcca7 · inbound

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs cites this paper.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:46.824918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:46.824918Z digest=sha256:ace7d013df34cc267bfcf7a92c2e4a7051268ee9bf4f957bc8b8795fbda5cdb8

Observation 4294a87b-37ca-4dd7-8684-3793d217ceef · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.416053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.416053Z digest=sha256:ce2d387b58ca1f8bbdd3543cdaacb9ba95f53211c75f4418070b8b4529cac7be

Observation ede358af-656a-4c9c-a204-21158acab1c5 · inbound

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond cites this paper.

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:25.023913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:25.023913Z digest=sha256:54da8a54a20f03a980afa43387f74a2fd0e6c5b7954e7d5297dcf379719b74a1

Observation c934db64-3e50-42a6-b795-2596f8f744e4 · inbound

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models cites this paper.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.543299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.543299Z digest=sha256:88a72439f0d16d3fac71826c0bb4218e9d1f0acbe32ce5588311a5616ce68bba

Observation bb2fa175-544f-4b44-8abd-a3127afaa569 · inbound

FACap: A Large-scale Fashion Dataset for Fine-grained Composed Image Retrieval cites this paper.

FACap: A Large-scale Fashion Dataset for Fine-grained Composed Image Retrieval TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:00.242206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:00.242206Z digest=sha256:67c8269256f89b17565c9060a6e529472069731ec84760542a1ae5d554d42920

Observation 0c424190-ddff-486c-a6ef-16aaf8b4e4ef · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.298449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.298449Z digest=sha256:766238932cb174fcc9e4a048a2e246ad102c0e2f35c65c4d0592f1ed38b1e09a

Observation 3243f840-cfe4-4a19-8f2d-3f79e8bfccb5 · inbound

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs cites this paper.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.589632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.589632Z digest=sha256:ea95fb6fe080dcf812606ecdf07053613a06bf961460a76bd1acde580d3968dd

Observation d3e74dd7-8293-4ec6-9204-f21aaa904a8f · inbound

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation cites this paper.

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:51.265156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:51.265156Z digest=sha256:951acd422ee26e82aa71ae81ba2bf2f35700847c1f1823e9fcd28596067c4027

Observation b04a3492-ffc0-4989-92c4-271b99e7bb8a · inbound

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces cites this paper.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.074772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.074772Z digest=sha256:d511fbd0b4d920dfdac6b08ff9c63beb937fc99773e0bf682ec950a294de22f8

Observation 1a3c9763-f343-4d6b-95d7-997527b62566 · inbound

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models cites this paper.

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T05:51:15.952781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:51:15.952781Z digest=sha256:330b512dddf2303d479d6ee2a63cb8c9fb774868a80dff6ea4f14d7b04557e15

Observation 9edf9c59-0678-425a-9243-243392fd084d · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.675460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.675460Z digest=sha256:03f3f6f57ffb9552530eb3f99df9d556557abc38a43faee7bf39dbe0843fdb5d

Observation 4b131883-97a4-4f14-97af-e2abeb8c3d79 · inbound

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors cites this paper.

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T13:05:54.422800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:05:54.422800Z digest=sha256:889f15bd25727ead16c3773b943de9417d0481d540627f7914468f03415867a8

Observation 8d937da5-6184-48f7-91e4-6eaf389f73b2 · inbound

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation cites this paper.

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:25:33.013951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T00:22:41.611893Z digest=sha256:f56e9e5d37e7b8301c968ac80c00d31e2fb0a36b83cb44e10c9aa17af45e09db

Observation 730fc652-2479-4f49-9878-86f56d5b5466 · inbound

UIPress: Bringing Optical Token Compression to UI-to-Code Generation cites this paper.

UIPress: Bringing Optical Token Compression to UI-to-Code Generation TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:00:59.210581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:21:32.024105Z digest=sha256:923655ef11aebe67a4448994654702766fc54ca1282b6948bcc51e9751f14de8

Observation de47aec6-9bef-40d4-9786-bfe0e2b88f70 · inbound

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding cites this paper.

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:15.864498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T08:13:42.526597Z digest=sha256:056a0dcc88bce812ed5a96ef2cc0a16ef0837715ffc9c72ff263da04d0b38473

Observation 2500fc84-b94d-418c-a669-5923dcef4800 · inbound

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs cites this paper.

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.619104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T05:41:04.184461Z digest=sha256:5e63994f2da36fd5a1bbbb258c45ae6063e7398a10d282af1daa51b5f4270545

Observation 8f953d7f-c32b-4170-910c-9dc353963321 · inbound

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning cites this paper.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.532712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.532712Z digest=sha256:d8ce17af8c9714e1623cfe4b66ed6a69f0815b215182fbd373defb44054fc888

Observation 029a34b8-1856-4299-8a32-a5378955ea5a · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.568021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.568021Z digest=sha256:4ecfac35d189c956dde3bcf4d8bcbab3db0122844662a63b36dc8d30fd9cbfc0

Observation a842a616-3f4c-4e81-9abb-82e1759158ba · inbound

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware cites this paper.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.805044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.805044Z digest=sha256:6954b559e678144abd5e4283d4ddb844b8c47ef405ced7b61a8688d3010af87b