Pith. sign in

Paper Citation Record · LEDGER

PuMer: Pruning and Merging Tokens for Efficient Vision Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2305.17530.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.17530 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:40:59.594437Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T07:34:22.099263Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2437dd00-81ac-4b3f-b64e-533ae57a91c1 · inbound

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference cites this paper.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference PuMer: Pruning and Merging Tokens for Efficient Vision Language Models

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.363405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:e2b5e8376682196cf7184330e1d6638c796f77cb62991428fdff90b5df82726b

Observation c2cc6f93-3ff4-4c25-a630-e9e490d7563b · inbound

A Survey on Mamba Architecture for Vision Applications cites this paper.

A Survey on Mamba Architecture for Vision Applications PuMer: Pruning and Merging Tokens for Efficient Vision Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T13:40:59.594437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:40:59.594437Z digest=sha256:de19f374a627c88793f2397279733468702f451ca3835e205d93c83bd06e5a30

Observation d84d369c-04f2-4233-bee2-acee828ac69c · inbound

Training-free Token Reduction for Vision Mamba cites this paper.

Training-free Token Reduction for Vision Mamba PuMer: Pruning and Merging Tokens for Efficient Vision Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:16:49.209222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:16:49.209222Z digest=sha256:594d4dff6dfe1f71e92d22f8829fb463ff000dfc7b3ee7983de2c77ae7e381b0

Observation b6d1382d-a1cc-41f7-a280-c75c178b5e50 · inbound

FastVGGT: Training-Free Acceleration of Visual Geometry Transformer cites this paper.

FastVGGT: Training-Free Acceleration of Visual Geometry Transformer PuMer: Pruning and Merging Tokens for Efficient Vision Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:36:06.918887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T23:36:06.870759Z digest=sha256:ef043e2aee94d82c14df385e09ac1730092ad4d3e34a6ce62fa6f0437d2c1790

Observation 1b766aaf-74ab-4420-bcf2-da20279652d5 · inbound

Accelerating Vision Transformers with Adaptive Patch Sizes cites this paper.

Accelerating Vision Transformers with Adaptive Patch Sizes PuMer: Pruning and Merging Tokens for Efficient Vision Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:42:24.457250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T05:41:48.429067Z digest=sha256:8d34bac97398546040ceebf166be63c00de0844fd862747b82f50a53c51abf90

Observation 746515ec-b6ac-4b36-b86b-3f498ab4b12c · inbound

Selective LoRA for Visual Tokens and Attention Heads cites this paper.

Selective LoRA for Visual Tokens and Attention Heads PuMer: Pruning and Merging Tokens for Efficient Vision Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:18:23.693969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T20:17:07.884892Z digest=sha256:85b77aa249334cb21582fa39ece5d33aac94dedbe40f365ea1757fcba49e5b58

Observation 8d94ad9c-c983-4730-a4e3-f4f66583d403 · inbound

Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies cites this paper.

Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies PuMer: Pruning and Merging Tokens for Efficient Vision Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:43:00.679014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T21:40:55.343169Z digest=sha256:5112d475593dd9820efa035e0fc0a0ce129231fd4d2bcc66a58ad7e916fb5213

Observation 5451c0bd-02b5-4667-a9ef-45d0cde960aa · inbound

Temporal Aware Pruning for Efficient Diffusion-based Video Generation cites this paper.

Temporal Aware Pruning for Efficient Diffusion-based Video Generation PuMer: Pruning and Merging Tokens for Efficient Vision Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.395585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T12:10:34.514379Z digest=sha256:9f8776f12ec55549cbc31c6b2e1e652803fd032d803888bc2586cfd64461f2eb

Observation 2fad2ed1-1040-40bb-9afb-b93c949fee7d · inbound

Temporal Aware Pruning for Efficient Diffusion-based Video Generation cites this paper.

Temporal Aware Pruning for Efficient Diffusion-based Video Generation PuMer: Pruning and Merging Tokens for Efficient Vision Language Models

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:24:47.156420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-22T10:21:33.661344Z digest=sha256:3907a5a8a562f7ce1760aec032dcfc33b145ca280148177e2c2d7382058363d9

Observation d065d733-4955-4e39-9b81-9062f9a82329 · inbound

State Machine Guided Multi-Relational Synthetic Data from Logs for Anomaly Detection cites this paper.

State Machine Guided Multi-Relational Synthetic Data from Logs for Anomaly Detection PuMer: Pruning and Merging Tokens for Efficient Vision Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:42:37.687954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T18:25:56.380366Z digest=sha256:a63036b81f02cf9556d7cb6fa6bbde472fd6c2a30f3cc68f57c89d05c3ee10cf

Observation 2887f1e4-bb4c-4ed7-a3f7-6017f8d40248 · inbound

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs cites this paper.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs PuMer: Pruning and Merging Tokens for Efficient Vision Language Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:34:22.100839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:9758d9e67fe282db2e41c058c739b5ece145da13ac67e4bd77eea5503396053f

Observation 9183615e-98c2-4696-83b1-0b1bdbca49c7 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception PuMer: Pruning and Merging Tokens for Efficient Vision Language Models

Reference 189

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:a38ff2fd1f1739022baa8875b150568293ff5a650b3533df28fd0f87d794f83b