Pith. sign in

Paper Citation Record · LEDGER

TokenPacker: Efficient Visual Projector for Multimodal LLM

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2407.02392.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.02392 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:47.521580Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:15:44.617819Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5ecbc9d4-5d8e-43c3-a3a6-0753f84636a3 · inbound

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step cites this paper.

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:35:25.953654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T11:35:25.894465Z digest=sha256:80b88d9ba2972fa98aa678b26dcf9c3e8739be19c9359f13af7ee46580c4ba5d

Observation 2ad7117d-d750-496d-bf95-cb4432bb11c7 · inbound

PixelThink: Towards Efficient Chain-of-Pixel Reasoning cites this paper.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.521580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:47.521580Z digest=sha256:dcaf5ad21d65853aaeb76a8f1ec93fc1418c2c7b965c47645a755c2aac5a0d86

Observation 9f5a24b9-6c20-4983-ad8f-d9e2bcdde671 · inbound

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models cites this paper.

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:04.930789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:09:04.930789Z digest=sha256:f9a27d7fb2f9fbc8716e69a6f50892f680705b46af649f8e12808e741f79e5e3

Observation 81818563-8738-439b-a096-f401ce88c5b9 · inbound

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding cites this paper.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.333184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.333184Z digest=sha256:e2c589669f8a0930b8b5ffd2615bc5f4ba09cca4c964ae6e4f8164a6c584c33b

Observation f1413e60-66b2-4474-8cc5-4e90d4930ee2 · inbound

Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings cites this paper.

Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:48.866029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:48.866029Z digest=sha256:c54f539dc845b4678b3ba3ae05a0fa616f03119bbab7af856188be2a18a2e95d

Observation e8315937-1c4c-4eed-a4d5-020faa7412bf · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:09.538863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:09.538863Z digest=sha256:bfe2a90009ee8c9e530ef454a97e71f5076bbd1829f07e1cb8ad788c8bd21dac

Observation 4adaa931-5207-495e-bbd1-7f0ca55cb3db · inbound

HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding cites this paper.

HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:47:07.697677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:47:07.697677Z digest=sha256:36a0776f71187efcd95872270ed6fbf82e867670462b4dcc9ff50c3b057de29a

Observation c62664f3-9f31-4c76-a0de-2a28e1bfcca7 · inbound

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs cites this paper.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:46.824918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:46.824918Z digest=sha256:ace7d013df34cc267bfcf7a92c2e4a7051268ee9bf4f957bc8b8795fbda5cdb8

Observation 4294a87b-37ca-4dd7-8684-3793d217ceef · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.416053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.416053Z digest=sha256:48d717008d2d6bb9af327532a17404755f7ea2c753697710073d1b6297f445d9

Observation ede358af-656a-4c9c-a204-21158acab1c5 · inbound

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond cites this paper.

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:25.023913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:25.023913Z digest=sha256:54da8a54a20f03a980afa43387f74a2fd0e6c5b7954e7d5297dcf379719b74a1

Observation c934db64-3e50-42a6-b795-2596f8f744e4 · inbound

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models cites this paper.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:29.543299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:29.543299Z digest=sha256:88a72439f0d16d3fac71826c0bb4218e9d1f0acbe32ce5588311a5616ce68bba

Observation bb2fa175-544f-4b44-8abd-a3127afaa569 · inbound

FACap: A Large-scale Fashion Dataset for Fine-grained Composed Image Retrieval cites this paper.

FACap: A Large-scale Fashion Dataset for Fine-grained Composed Image Retrieval TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:00.242206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:00.242206Z digest=sha256:67c8269256f89b17565c9060a6e529472069731ec84760542a1ae5d554d42920

Observation 0c424190-ddff-486c-a6ef-16aaf8b4e4ef · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.298449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.298449Z digest=sha256:766238932cb174fcc9e4a048a2e246ad102c0e2f35c65c4d0592f1ed38b1e09a

Observation 3243f840-cfe4-4a19-8f2d-3f79e8bfccb5 · inbound

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs cites this paper.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.589632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.589632Z digest=sha256:6f8d57af11c7d8f6cf4e0c916baff1b6ba561ddcfebf60d42fb8019da51ad980

Observation d3e74dd7-8293-4ec6-9204-f21aaa904a8f · inbound

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation cites this paper.

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:51.265156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:51.265156Z digest=sha256:8a990dfa33e82231563b2977314c473f97be376ce88ae123d1eaed1f2297d8d1

Observation b04a3492-ffc0-4989-92c4-271b99e7bb8a · inbound

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces cites this paper.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.074772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.074772Z digest=sha256:d511fbd0b4d920dfdac6b08ff9c63beb937fc99773e0bf682ec950a294de22f8

Observation 1a3c9763-f343-4d6b-95d7-997527b62566 · inbound

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models cites this paper.

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T05:51:15.952781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:51:15.952781Z digest=sha256:330b512dddf2303d479d6ee2a63cb8c9fb774868a80dff6ea4f14d7b04557e15

Observation 9edf9c59-0678-425a-9243-243392fd084d · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.675460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.675460Z digest=sha256:03f3f6f57ffb9552530eb3f99df9d556557abc38a43faee7bf39dbe0843fdb5d

Observation 4b131883-97a4-4f14-97af-e2abeb8c3d79 · inbound

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors cites this paper.

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T13:05:54.422800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:05:54.422800Z digest=sha256:889f15bd25727ead16c3773b943de9417d0481d540627f7914468f03415867a8

Observation 8d937da5-6184-48f7-91e4-6eaf389f73b2 · inbound

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation cites this paper.

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:25:33.013951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T00:22:41.611893Z digest=sha256:7ae040fc8ffc17b488ed56137787c075077f8787a27f4bdf95eae7500fe34e27

Observation 730fc652-2479-4f49-9878-86f56d5b5466 · inbound

UIPress: Bringing Optical Token Compression to UI-to-Code Generation cites this paper.

UIPress: Bringing Optical Token Compression to UI-to-Code Generation TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:00:59.210581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:21:32.024105Z digest=sha256:9b3a80ece71c8f5fcbf1d82c55aedf00cae2c1999c6e1bb70977779fd6668000

Observation de47aec6-9bef-40d4-9786-bfe0e2b88f70 · inbound

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding cites this paper.

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:15.864498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T08:13:42.526597Z digest=sha256:517248f067f6e16494e2a08e38218c71cc9031118e0888f12c40949ac51f2a43

Observation 2500fc84-b94d-418c-a669-5923dcef4800 · inbound

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs cites this paper.

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.619104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T05:41:04.184461Z digest=sha256:6e4648fe30106ad31fdfdd3794295d9517f56f96daa9b28bc6ed02cb803fde92

Observation 8f953d7f-c32b-4170-910c-9dc353963321 · inbound

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning cites this paper.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.532712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.532712Z digest=sha256:eaf2cb432aa6aecab3dfb686dee2a8ac31b6b2179b5f4986a3d7bcaab33e383d

Observation 029a34b8-1856-4299-8a32-a5378955ea5a · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.568021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.568021Z digest=sha256:4ecfac35d189c956dde3bcf4d8bcbab3db0122844662a63b36dc8d30fd9cbfc0

Observation a842a616-3f4c-4e81-9abb-82e1759158ba · inbound

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware cites this paper.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.805044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.805044Z digest=sha256:6954b559e678144abd5e4283d4ddb844b8c47ef405ced7b61a8688d3010af87b