Pith. sign in

Paper Citation Record · LEDGER

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models

As of 12 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2607.13500.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13500 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:03:30.238634Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f4ab3642-ce73-454d-b2c0-611e1b14a6c4 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:24.436267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:24.436267Z digest=sha256:18a67d77e4ea1aaf39f56b720ce4745c0ffe05bd92dd82f850c5f74ea65a6e2e

Observation c9847a69-a0ba-41e3-afb5-90c2ec4a0642 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:24.536860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:24.536860Z digest=sha256:a9f0bf825ac558780f4597fdb78f025b69af08f1a894517ea4d8488a94f48186

Observation 226cf9b0-c893-49a1-bc4f-2e0510e1024a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:24.689482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:24.689482Z digest=sha256:828a32aa6011f1a01ebbf5a153c86297f7261aa16a75dcf685db429cbac18cd6

Observation ea241e75-2bb5-4ef9-b833-a0bc3a062b41 · outbound

This paper cites BLIP- 2:bootstrapping language-image pre-training with frozen 11 image encoders and large language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models BLIP- 2:bootstrapping language-image pre-training with frozen 11 image encoders and large language models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:24.814141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:24.814141Z digest=sha256:37e9dbb7a4fc6f2bce2c0c10dbf164d257e9db1490ee358f4dae8d94215258e9

Observation 9b2ade47-0381-4662-9bb0-0a2d01c718ef · outbound

This paper cites Cost-efficient and secure federated learning for edge computing,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Cost-efficient and secure federated learning for edge computing,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:24.897278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:24.897278Z digest=sha256:c6bea6f522db08f497fe2f1c8b628039353ea8cf936cc82b997b59c87a064c39

Observation 0dbdf4ac-b3dd-40f0-9467-1989ec35030a · outbound

This paper cites To- wards online privacy-preserving computation offloading in mobile edge computing,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models To- wards online privacy-preserving computation offloading in mobile edge computing,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:25.037156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:25.037156Z digest=sha256:c170d09718e702c7ed445dc4b7ba8c159edb8e9bce4ed76a276c6fe00e9f375c

Observation c110d243-9905-40b3-a029-e4f09786feb6 · outbound

This paper cites AFLoRA: Adaptive Federated Fine-Tuning of Large Language Models with Resource-Aware Low-Rank Adaption.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models AFLoRA: Adaptive Federated Fine-Tuning of Large Language Models with Resource-Aware Low-Rank Adaption

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:25.152278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:25.152278Z digest=sha256:9a9d2f09b771acc9714d3dda6bc97adea128b893c121a91c06663c46a48893cc

Observation 54509875-4356-453a-92fd-965ad18dd9bd · outbound

This paper cites Towards efficient edge learning for large models in heterogeneous resource-limited environments,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Towards efficient edge learning for large models in heterogeneous resource-limited environments,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:25.325447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:25.325447Z digest=sha256:f7ed5f21452ad7fa3f8050ed37d12cc7a1897cbbd2565117e020cc851757b4dd

Observation 62585956-b6f2-4224-b033-c57ae6dbf983 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:25.434759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:25.434759Z digest=sha256:99250f377a023354431e21e76d5ee160755aabca475c5e262029517bc71dc112

Observation 1e4c373d-8840-48a8-b27f-9bdcb9d49091 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Learning transferable visual models from natural language supervision,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:25.501494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:25.501494Z digest=sha256:9aa8a0a66a648a47fdf14ccffd25b2fd62e73709578e763350c543d7e149d110

Observation b0e52ff1-b723-42a7-a6ba-b6613a029b58 · outbound

This paper cites Sigmoid loss for language image pre-training,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Sigmoid loss for language image pre-training,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:25.611352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:25.611352Z digest=sha256:2b0c44fef367bc2cbf4152521bc38eb2a55177b07ac84726f64b56480363c732

Observation b3c0b6c8-aa10-4928-92c3-7aefb0c0232c · outbound

This paper cites Tap-vits: Task-adaptive pruning for on- device deployment of vision transformers,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Tap-vits: Task-adaptive pruning for on- device deployment of vision transformers,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:25.794944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:25.794944Z digest=sha256:772277b75fc857e6c575d60fc2ed52d13185d34f63e462447b0a9275593b399b

Observation 4d088e23-6195-4a16-bcc8-5d894330ded2 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision- language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision- language models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:25.924755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:25.924755Z digest=sha256:2a20f8a53e324ce521c91590a42c6c33d72fdace3676f48cd45198f15bb25d2b

Observation f5b2543d-5b97-41a7-82d5-84b2e311bc4e · outbound

This paper cites VScan: Rethinking visual token reduction for efficient large vision-language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models VScan: Rethinking visual token reduction for efficient large vision-language models,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:26.059229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:26.059229Z digest=sha256:7945ad1a9b08bb56257f8518f03b4fdeb9e1d27e4df2af97c5db0d27099f3d46

Observation 064b0bc9-7083-443c-8844-053fe9866142 · outbound

This paper cites Adaptinfer: Adaptive token pruning for vision-language model inference with dynamical text guidance,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Adaptinfer: Adaptive token pruning for vision-language model inference with dynamical text guidance,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:26.144959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:26.144959Z digest=sha256:bff1dc364940886c2b303802aee442078de24b8184763b36deca7fdbd47f90dc

Observation 7fc6fe26-451d-44fc-84b4-94e73926579c · outbound

This paper cites Variation-aware vision token dropping for faster large vision-language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Variation-aware vision token dropping for faster large vision-language models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:26.344745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:26.344745Z digest=sha256:519741a94d5dc0a58bf9ed2c63aa1979c0da5b42c2e049a05c64de019242f974

Observation b7a58740-0054-4df7-bbdd-43b100638dd8 · outbound

This paper cites Sparsevlm: Visual token sparsification for efficient vision-language model inference,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Sparsevlm: Visual token sparsification for efficient vision-language model inference,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:26.444741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:26.444741Z digest=sha256:0bd16352befc88d23e4564b0594b9e186958160e3a3fd30fdd4c3f518bb3296d

Observation 6e3d6150-0858-4287-b4a9-ec64c80658be · outbound

This paper cites Pyramiddrop: Accelerating your large vision-language models via pyra- mid visual redundancy reduction,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Pyramiddrop: Accelerating your large vision-language models via pyra- mid visual redundancy reduction,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:26.552698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:26.552698Z digest=sha256:1239b45ddbae8d24d56831a5382cf795ffa6497a76fab44b603b4f11b88e7f15

Observation 13493d25-1882-4ac2-9c22-5251d778b29b · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:26.664752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:26.664752Z digest=sha256:c07e116bab291c9312eddf96f50e095af693c136bee481c3cbaf94fa70fcdfab

Observation d13bbfb4-1f49-43db-b16c-80d33edd3730 · outbound

This paper cites LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:26.782904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:26.782904Z digest=sha256:a6c592168b6a3e52666cf4d7376a61acb5a4cc3674b564618914f612e69eb0da

Observation 0ff3eba6-3f51-43f7-a7fb-97b6a2b91979 · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Visionzip: Longer is better but not necessary in vision language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:26.870445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:26.870445Z digest=sha256:76e45ac9ffa75f472580751876c4de0bd0fa6af54a8f1a125b820887eec49d41

Observation b084ba45-397d-4a07-89af-fc30c640744a · outbound

This paper cites Beyond text-visual attention: Exploiting visual cues for effective token pruning in vlms,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Beyond text-visual attention: Exploiting visual cues for effective token pruning in vlms,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:26.948681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:26.948681Z digest=sha256:b571c107e3b49d6b2879f4b95f7cb9099775cdff8c8b6b50b904fbbb20b60840

Observation 546ee96d-be61-4ad7-84d2-45d3f3cbd165 · outbound

This paper cites Flashat- tention: Fast and memory-efficient exact attention with io-awareness,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Flashat- tention: Fast and memory-efficient exact attention with io-awareness,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:27.092271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:27.092271Z digest=sha256:2921a1427b9a9f863a37ab49e5182c9b658efd73690cc8f870bf5fb5470326e5

Observation 2deb5366-dbe8-4b57-bbb0-61b467999fb5 · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Flashattention-2: Faster attention with better parallelism and work partitioning,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:27.234826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:27.234826Z digest=sha256:f04a57cb5f126b165c9b4ca1dca875648585d689ffefa330586be32403522706

Observation cc87243f-39eb-46ac-bcda-da8422f6b7ce · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Flashattention-3: Fast and accurate attention with asynchrony and low-precision,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:27.403994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:27.403994Z digest=sha256:e2e8a637deea5e684fbb1a426aba3e0a6319579881dc2734502dfe5644bf66e6

Observation 59915c0f-7492-4139-8a10-0ad3127b019b · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Efficient memory management for large language model serving with pagedattention,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:27.538342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:27.538342Z digest=sha256:f47667db5489dbe50d47f72e4830c54e5545248b0895c316d1f9b15d924205dd

Observation 9774a8cb-bf6f-4697-814e-07c4097cb51e · outbound

This paper cites Similarity-Aware Token Pruning: Your VLM but Faster.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Similarity-Aware Token Pruning: Your VLM but Faster

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:27.630859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:27.630859Z digest=sha256:f2c21e51517b5be1cf924ce881f11bec3143a78afa2f026e5f9be577cfcea1af

Observation 22517f44-5ec4-418c-aa52-cf96dcdb49a2 · outbound

This paper cites DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:27.729468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:27.729468Z digest=sha256:6efd5eccce7a27bd642c015b83579b4ffdc4cf99cd7eaf174430b0ad81f818ec

Observation 3d64b4bf-0c8f-4592-af83-ac3c21952221 · outbound

This paper cites Holitom: Holistic token merging for fast video large language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Holitom: Holistic token merging for fast video large language models,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:27.825305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:27.825305Z digest=sha256:bfb7aa6bbd6feb5d7b445ecbd59c164291167e38cbaebbd910b5b3336b4bd120

Observation 5205495b-aa63-4cf8-b571-2f04b3603e2c · outbound

This paper cites Dycoke: Dynamic compression of tokens for fast video large language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Dycoke: Dynamic compression of tokens for fast video large language models,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:27.926783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:27.926783Z digest=sha256:2818abcd0477deba7859c069205b7c7ccc9a53e7eebc6efe2a9226bc5467de81

Observation 87784a5b-205f-46f8-b004-c1470cc9e687 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and- language tasks,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and- language tasks,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:28.013830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:28.013830Z digest=sha256:f28168a80826bed95cbdc250e62f40f748d8f6d61f70f48939c920a1b98f7720

Observation be0da18c-1d61-449f-9964-be0b68b85197 · outbound

This paper cites Uniter: Universal image-text representation learning,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Uniter: Universal image-text representation learning,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:28.121689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:28.121689Z digest=sha256:2edf5844f6c85fa6bb31837abec3da7ceb5c2fa75c4c12b4cafcf2832b182368

Observation d8089c62-c843-492d-ad37-af8967cc2be0 · outbound

This paper cites Vflowopt: A token pruning framework for lmms with visual information flow-guided optimization,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Vflowopt: A token pruning framework for lmms with visual information flow-guided optimization,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:28.226288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:28.226288Z digest=sha256:5620385979120bc9c3bcf57f11a2c866f8954ddea908bae064c668595ce7da56

Observation 6dd0f4b1-427a-4618-b40c-39e40a799b4a · outbound

This paper cites Hidrop: Hierarchical vision token reduction in mllms via late injection, concave pyramid pruning, and early exit,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Hidrop: Hierarchical vision token reduction in mllms via late injection, concave pyramid pruning, and early exit,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:28.371847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:28.371847Z digest=sha256:7372a4a15b4a325f09799c62505f743f413fe9a696424326dee1c22d4be8cee1

Observation 4e8aa322-7b15-44e8-bcba-029479783c31 · outbound

This paper cites Todre: Visual token pruning via diversity and task awareness for efficient large vision- language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Todre: Visual token pruning via diversity and task awareness for efficient large vision- language models,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:28.468713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:28.468713Z digest=sha256:e47a61bfa5aaa012193159524a56d04ca5706f18afa4c19a7deb7c0ca45f77de

Observation c635a1e2-7d3a-4d3b-8cb3-6c77aa6b9dab · outbound

This paper cites Tamp: Token-adaptive layerwise pruning in multimodal large language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Tamp: Token-adaptive layerwise pruning in multimodal large language models,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:28.594759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:28.594759Z digest=sha256:b3728e5ea4dfce0a9355ddb7a95e9e29f2071bc1ed4984093717005467e86d12

Observation c19994dd-9808-4d3f-8bd0-6836fccd94db · outbound

This paper cites Swiftvlm: Efficient vision-language model inference via cross-layer token bypass,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Swiftvlm: Efficient vision-language model inference via cross-layer token bypass,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:28.694745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:28.694745Z digest=sha256:7998ea387caa77eb524ce2d24fe464fce7f0112b7efcf4a80ef8d0b5cea50a2c

Observation 132bf8bb-d0dd-40d9-96a0-259f560bd421 · outbound

This paper cites Fit and prune: Fast and training-free visual token pruning for multi-modal large language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Fit and prune: Fast and training-free visual token pruning for multi-modal large language models,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:28.821960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:28.821960Z digest=sha256:899fb942320cd695d75420914ae941eaedaab0fc3e2b4af490cc8cdf6b189f42

Observation f5b61b77-33b6-4232-bd3b-dfeaa1c807e6 · outbound

This paper cites Llava-mini: Efficient image and video large multimodal models with one vision token,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Llava-mini: Efficient image and video large multimodal models with one vision token,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:28.915880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:28.915880Z digest=sha256:2f101b8d43cdb43cf6f59b90cfc628fa517e7c09e32b59c551d9b0786a2fbe5d

Observation cda51eae-d9b7-4f65-9f58-6e8067fb0f7a · outbound

This paper cites Lvpruning: An effective yet simple language- guided vision token pruning approach for multi-modal large language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Lvpruning: An effective yet simple language- guided vision token pruning approach for multi-modal large language models,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:29.080528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:29.080528Z digest=sha256:8fe85929c2487bd5e35f832551c4bea56c9dd83d5b38928002b36f7de3d583f0

Observation 1edaa9e9-2c2a-4a06-bcef-5a0127102bd0 · outbound

This paper cites Matryoshka multimodal models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Matryoshka multimodal models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:29.165018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:29.165018Z digest=sha256:c83a8d23be486097f3329486ee641d54ed44e7df56c18c0e42c99d7ea82e0c3e

Observation a558b01d-b723-454a-b8f6-6bbf7e44ee96 · outbound

This paper cites Matryoshka query transformer for large vision-language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Matryoshka query transformer for large vision-language models,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:29.310156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:29.310156Z digest=sha256:07ec3dd8541e6d034622966226758551b0fbf13c8658211ed747a963ee24d0cc

Observation 089ab825-792e-428a-9c62-f70ae9564a87 · outbound

This paper cites Improved baselines with visual instruction tuning,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Improved baselines with visual instruction tuning,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:29.446041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:29.446041Z digest=sha256:47c378909d4403e8d3462db5b0d92bdff0be6c5caf6f5a8036925a3de5d1a30d

Observation 3626fb5f-c1ba-4f9a-8255-a056036adf9c · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:29.497826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:29.497826Z digest=sha256:fea114cbca120bff7cfb5a19ed7ab1cfe4beef25cfe98d3bbb1940fbce72b2cb

Observation d2bec4f0-55dc-4e30-8613-937675819682 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Mme: A comprehensive evaluation benchmark for multimodal large language models,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:29.612676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:29.612676Z digest=sha256:cbd049c4b9b45cfea9670294dc7786fed9069e730b6f4f4cdb2d2aa8c1d1f565

Observation b00ddbee-7c28-4e7f-ae45-562ce7929f69 · outbound

This paper cites Towards vqa models that can read,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Towards vqa models that can read,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:29.779055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:29.779055Z digest=sha256:1bbba638edc70613211d777b919a48d6e336ab36de0c9e5ba848b22e9dd71d78

Observation 773bc19c-4e06-4a80-a624-ee3a1a702f9c · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:29.912039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:29.912039Z digest=sha256:83bb52dd3f7cfe25423a67d2dd1c2b44806160fd383f269279ccdc4427f8180f

Observation ddf231ed-aef2-4a71-b883-5cee9fa5a70d · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:30.032616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:30.032616Z digest=sha256:89a1e6d5d6d372101cc3f269d581a35ff86a933187791e0846ef8e1c37254fc3

Observation b11c6af6-3425-4468-a3a1-a874fad95616 · outbound

This paper cites Qwen2.5-vl,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Qwen2.5-vl,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:30.141140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:30.141140Z digest=sha256:4fd622ed2cce4da3914d7f59aca53f770174236edaa30a2ddf691b59110f31c3

Observation 71ae835d-ec1b-4842-9824-74af18c345bd · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library,.

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Pytorch: An imperative style, high-performance deep learning library,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T05:03:30.238634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:03:30.238634Z digest=sha256:f54042300a0d3394330c387442b3dff8d2a395aba71184cad4e066c69add528e

Pith citing papers

No inbound Pith citation observations are available.