Pith. sign in

Paper Citation Record · LEDGER

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration

As of 16 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 3 inbound Pith citation observations for arXiv:2506.05709.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05709 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:29:09.281677Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:40:36.389460Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T13:20:26.673617Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0888780a-745f-43c7-a35a-fbecf752f310 · outbound

This paper cites Token Merging: Your ViT But Faster.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Token Merging: Your ViT But Faster

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:04.281431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:04.281431Z digest=sha256:b667242dc9ccb9d467f564238fc7449200b477f349f568e3f185c9d035dd46f2

Observation 5cab2b06-3244-4ce3-831a-2e70ad67fb4a · outbound

This paper cites Crossvit: Cross-attention multi-scale vision transformer for image classification.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Crossvit: Cross-attention multi-scale vision transformer for image classification

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:04.388730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:04.388730Z digest=sha256:64e0f1807ebffae63cee948917d6a6179f015f3a13245bd488ca7c42979f0e20

Observation 0b1abe6f-8215-4308-be29-9f31d4057c79 · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration The cityscapes dataset for semantic urban scene understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:04.502992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:04.502992Z digest=sha256:aa5f52daaab0fd97b37b2f377fe83965243725dcf07db33c845653d020e2b552

Observation 4fb8e953-7fe6-4d69-868f-042896b35c86 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Imagenet: A large-scale hierarchical image database

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:04.652087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:04.652087Z digest=sha256:bd863159589a56ae8a29efc35c9505a60bdc0a01161ad6a549a4e9b7cfc4dbed

Observation e9372569-a820-4c8b-b1ed-e90341c7b665 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:04.745969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:04.745969Z digest=sha256:5fa165b5794183e208ae7d9f5c6525e6dbb7b4a4cf7f4ae375a1a6184b8ff9ca

Observation f5b48f0c-07db-4424-98a3-57a965da433f · outbound

This paper cites Adaptive token sampling for efficient vision transformers.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Adaptive token sampling for efficient vision transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.711924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:04.879291Z digest=sha256:2a67569034e919574962816d90484f4987adee6283094719150abb0a43d4c539

Observation a64d3003-f149-462d-8504-445086b060fb · outbound

This paper cites Dynamic Channel Pruning: Feature Boosting and Suppression.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Dynamic Channel Pruning: Feature Boosting and Suppression

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:04.957294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:04.957294Z digest=sha256:69d395b1e278892be46f0bb24321d39d6df332499efd3d7033544bb00f111bca

Observation ef0f1ead-73e8-4bac-b6ee-714e548b1e29 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:05.047412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:05.047412Z digest=sha256:96a4958193dfa754ab86c81486f9671146a1c7197e5f458cfa1c2dd0fe9a3377

Observation b28c7a65-1c10-4b2c-8014-f04eef38c082 · outbound

This paper cites Dynamic neural networks: A sur- vey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11):7436–7456, 2021.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Dynamic neural networks: A sur- vey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11):7436–7456, 2021

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.681292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:05.135319Z digest=sha256:0189c935a7533c91f80cdd3e8205de4a48c2ef899ff6f728af6e8ffe98e5dad6

Observation 46435033-fb1a-4eb3-ad6e-fba69d2779bf · outbound

This paper cites Masked autoencoders are scalable vision learners.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Masked autoencoders are scalable vision learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:05.210527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:05.210527Z digest=sha256:987f66a9c257ded154601d0cad7ae6a168484c601065c61d724438870ac218ca

Observation b56a5213-dfa2-4563-9664-36861b6d4660 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Distilling the Knowledge in a Neural Network

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:05.342081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:05.342081Z digest=sha256:17a2b71a4171ff752d92a6a5286b0a3c17fa4afec7edc466d38df3171a2a8ff6

Observation bc203aa9-11cd-4c25-9936-ab7f1167166e · outbound

This paper cites All tokens matter: Token labeling for training better vision transform- ers.Advances in neural information processing systems, 34: 18590–18602, 2021.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration All tokens matter: Token labeling for training better vision transform- ers.Advances in neural information processing systems, 34: 18590–18602, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.651394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:05.430042Z digest=sha256:e77e031644c490db072ef6dfe0ae049052ad37fc01279bad6bc3c538872df841

Observation 0db0ec9d-8003-4d08-9186-e582ca2adfaa · outbound

This paper cites Spvit: Enabling faster vision transformers via latency-aware soft token pruning.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Spvit: Enabling faster vision transformers via latency-aware soft token pruning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.624831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:05.506428Z digest=sha256:9fe7818828a27c1e6ace1e9bb9b0b1f6db4021626fbed487c83eaf68cbb4ff31

Observation a38b61e1-ebe5-4e71-a9c9-425a8547639a · outbound

This paper cites Peeling the onion: Hierarchical reduction of data redundancy for efficient vision transformer training.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Peeling the onion: Hierarchical reduction of data redundancy for efficient vision transformer training

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.606625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:05.585257Z digest=sha256:810b35335ece1adcf3e696aae66bbed4bc73f4d9d922dec49cf4a7167ec41d1d

Observation 45dc1c5f-6e8c-4412-9a90-13b27f30a900 · outbound

This paper cites Token reduction should go beyond effi- ciency in generative models–from vision, language to mul- timodality.arXiv preprint arXiv:2505.18227, 2025.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Token reduction should go beyond effi- ciency in generative models–from vision, language to mul- timodality.arXiv preprint arXiv:2505.18227, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:05.660989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:05.660989Z digest=sha256:791c2afce911e0f6de8e38ed38fa31453f1229d62755ea69ea71e73c3a752389

Observation 7745153f-5938-43a9-9bcb-ddd347d7c914 · outbound

This paper cites Mpvit: Multi-path vision transformer for dense prediction.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Mpvit: Multi-path vision transformer for dense prediction

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.580185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:05.744711Z digest=sha256:bc70e6423915905e83ae1d735722064caf6dc9c38524e133142cd8ae3c2e29d4

Observation ebd511db-f48c-457c-a81e-9910295f8dfb · outbound

This paper cites Vidtome: Video token merging for zero-shot video editing.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Vidtome: Video token merging for zero-shot video editing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.554998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:05.823630Z digest=sha256:cbf5e3949c39f2bceb6c52a9550f85310a8273e0a3b64e10f48b7a9f59966120

Observation 35781d02-da4d-4d47-b89e-e82f19139409 · outbound

This paper cites Exploring plain vision transformer backbones for object de- tection.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Exploring plain vision transformer backbones for object de- tection

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.527097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:05.934640Z digest=sha256:9e3012c561cd7ac795d24510a2ed558b7e9cb60aa8df850f35f97546fe94ec74

Observation 6f97ca6a-6c2a-441f-bf9b-056cfbe1cec5 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Evaluating object hallucination in large vision-language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.499169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:06.014456Z digest=sha256:12f4fc153a1782a88950a23d72e4edd25d2c508bcd54dc80f1eb5d3798a59740

Observation 3187dbc6-bf1c-4d47-b63b-1d61bb3e8225 · outbound

This paper cites Expediting large-scale vision transformer for dense predic- tion without fine-tuning.Advances in Neural Information Processing Systems, 35:35462–35477, 2022.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Expediting large-scale vision transformer for dense predic- tion without fine-tuning.Advances in Neural Information Processing Systems, 35:35462–35477, 2022

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.476808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:06.111149Z digest=sha256:e9945431786fa6764fdcbaf73fe25caa13323f8862e2982ffa2f9b2580a50dcf

Observation 4ea22053-9204-4db7-8171-e2f7f4ebb9c3 · outbound

This paper cites Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:06.229026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:06.229026Z digest=sha256:c5bd0245145055dbb4d428481e709f98a34a744eb1ad2d2cf6984f5a3845f310

Observation 12b87abb-f5f7-4210-b8ce-fdacd94047db · outbound

This paper cites Microsoft coco: Common objects in context.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Microsoft coco: Common objects in context

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:06.305116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:06.305116Z digest=sha256:f404a127effb1d62ab9bbff7cd3ff7a30857cf9e6a8f913a0ff367d92df03191

Observation 43770335-d3c0-455d-ad52-55898dd40e3d · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:06.392937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:06.392937Z digest=sha256:cc2a23b78a74c279647106b304abedb756d1bd0c2d3ef0d319ead7c245616579

Observation 860ec1a7-aee1-403f-8680-eb4b3b6ebb18 · outbound

This paper cites Adaptive Sparse ViT: Towards Learnable Adaptive Token Pruning by Fully Exploiting Self-Attention.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Adaptive Sparse ViT: Towards Learnable Adaptive Token Pruning by Fully Exploiting Self-Attention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:06.482067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:06.482067Z digest=sha256:1c743496ffafb3fc1469d7fcb60bc7d3fffbee8a56d83e530b72b3c72889444f

Observation 33436fba-e7b2-4490-ada3-601545fe07d2 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Swin transformer: Hierarchical vision transformer using shifted windows

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:06.552259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:06.552259Z digest=sha256:591ecfd645a8c41465510624feb61079748a3190d832b2a150d28cdc06478716

Observation 979a44d9-6da2-46b2-9c21-3b0c8a1f665a · outbound

This paper cites Beyond attentive tokens: Incorporating to- ken importance and diversity for efficient vision transform- ers.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Beyond attentive tokens: Incorporating to- ken importance and diversity for efficient vision transform- ers

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.428964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:06.636831Z digest=sha256:0c0fd8317006acc8904f9badd1ec1552f9e77d0c5bd11f3a301ab178d5e897fa

Observation 8009a827-c332-4af3-857d-f6ca9329f6b4 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:06.746497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:06.746497Z digest=sha256:97320e10bade71d2c46978147ef453dc2b3b5c307237bc9d1259460286ef08ed

Observation 2c5cf6f9-377d-48bf-9ce6-902d88ad763e · outbound

This paper cites Importance estimation for neural net- work pruning.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Importance estimation for neural net- work pruning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.391781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:06.834295Z digest=sha256:cb9712d539f163bba4096805678f75d400d672281c8d13ce907c059a9927a278

Observation 389c33d8-09c4-43d5-8a3d-b274e8ee555d · outbound

This paper cites Improving language understanding by gen- erative pre-training.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Improving language understanding by gen- erative pre-training

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:06.910131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:06.910131Z digest=sha256:d36e588e6ae98a2df4af79b251ac1c07e7f57db76ac7aec16a77a74ecb081e71

Observation be14cf72-3f22-4df1-aa32-85e1491bf8af · outbound

This paper cites Vi- sion transformers for dense prediction.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Vi- sion transformers for dense prediction

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.130693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:06.981966Z digest=sha256:6f09488639ee4b6bae87834db1708ddf5b566ed630b6cbd32cd4150b159719b4

Observation bd498842-e345-4559-80c2-2c70f62f00f0 · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.Advances in neural information processing systems, 34:13937–13949,.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Dynamicvit: Efficient vision transformers with dynamic token sparsification.Advances in neural information processing systems, 34:13937–13949,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:07.068786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:07.068786Z digest=sha256:53e7bd0a723c0dddaf244d5c19e5b28925e565614f56ee9b144f2e92f286b768

Observation c4914461-08cc-4b41-aa12-1558c457f3b5 · outbound

This paper cites Beyond fixa- tion: Dynamic window visual transformer.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Beyond fixa- tion: Dynamic window visual transformer

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:42.989923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:07.153073Z digest=sha256:bd59f3f1355045ae607bfa6a460a4af37235e5e58678e9b0f66c69593380465f

Observation c15a2e28-18e9-45f3-b269-8df503a92712 · outbound

This paper cites Tokenlearner: Adaptive space-time tokenization for videos.Advances in Neural In- formation Processing Systems, 34:12786–12797, 2021.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Tokenlearner: Adaptive space-time tokenization for videos.Advances in Neural In- formation Processing Systems, 34:12786–12797, 2021

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:42.845527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:07.225649Z digest=sha256:fb896571a6e39849088f17cc2dd3d9cd9192a8a9bc1fb06e689922a8bf7d907e

Observation 599cf5bd-dd11-4724-8683-bef66c8ba766 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.arXiv preprint arXiv:2403.15388,.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Llava-prumerge: Adaptive token reduction for efficient large multimodal models.arXiv preprint arXiv:2403.15388,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:07.294589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:07.294589Z digest=sha256:1c479dce1b42f63f98eb60411ee87e0bedf0c54d1198c704492fa9d9403029f1

Observation 0c617dbe-ca68-4420-8bfa-56e394ba68d8 · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Indoor segmentation and support inference from rgbd images

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:07.373418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:07.373418Z digest=sha256:a649ac60f0f7509df6c49f584e84e8e388f9a2867831737d149413e037d0d135

Observation f38efc45-34f6-4c89-9451-0a85d54ac8b4 · outbound

This paper cites Towards vqa models that can read.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Towards vqa models that can read

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:07.438068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:07.438068Z digest=sha256:2faeb68ecc102bd54970853551e1ec4d012b471c3e7edb355e3c481e2549f0b0

Observation dc0e95be-ffb6-4161-bdbf-bf987a02c641 · outbound

This paper cites How to train your vit? data, augmentation, and regularization in vision transformers.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration How to train your vit? data, augmentation, and regularization in vision transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:42.738363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:07.536108Z digest=sha256:f49a3da6f67cce5042fa30cf1f5bc95fc312d070d8beb1e61448f89b08c0d7c6

Observation 180647af-62ad-4a1c-a0c8-5fd678508398 · outbound

This paper cites Segmenter: Transformer for semantic segmenta- tion.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Segmenter: Transformer for semantic segmenta- tion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:07.626801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:07.626801Z digest=sha256:c0fdfd01c438a859461466c1f2c443c76c159db3691ba5a8621e181edbb3b50a

Observation 3bf620c2-ed9e-41c9-9bb4-b3d7605a99f3 · outbound

This paper cites Chip: Channel independence- based pruning for compact neural networks.Advances in Neural Information Processing Systems, 34:24604–24616,.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Chip: Channel independence- based pruning for compact neural networks.Advances in Neural Information Processing Systems, 34:24604–24616,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:07.694796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:07.694796Z digest=sha256:8106f9c19cce2cc257edba192cef53e8462b80703b81fe876bc639c615d3cf8d

Observation 665c0120-0de5-4f2f-b165-53a8a84b86ed · outbound

This paper cites Revisiting unreasonable effectiveness of data in deep learning era.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Revisiting unreasonable effectiveness of data in deep learning era

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:42.560510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:07.784186Z digest=sha256:acd6b4061dad246caebeea23f245e3524e2dc35a7c506b882fd4ee23b8acc6d1

Observation 7642238b-05ad-4b54-95bd-5a92d8873de0 · outbound

This paper cites Patch slimming for ef- ficient vision transformers.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Patch slimming for ef- ficient vision transformers

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:42.468333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:07.843474Z digest=sha256:c17c78f8e3608ebaa1dcc053c18d9cc902586b37e5c98cc9796ae7650bb5c620

Observation 63b54ada-a354-421e-8569-7910f6534099 · outbound

This paper cites Training data-efficient image transformers & distillation through at- tention.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Training data-efficient image transformers & distillation through at- tention

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:11.049320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:07.913513Z digest=sha256:c420002332710ba7e7dd054303d61395d12454010ad4e35190270bd265b04c41

Observation e790d5cb-ccf1-4974-9c54-65424d8265f7 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:07.974529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:07.974529Z digest=sha256:a5d8c34305adb9a352bd0fcd6df6f796dcc65550a21fb1ce53734265dae82a1f

Observation 6265c65b-848b-463c-846e-00f2bf5bfa08 · outbound

This paper cites Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition.Advances in Neural Information Processing Systems, 34:11960–11973,.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition.Advances in Neural Information Processing Systems, 34:11960–11973,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:10.912369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:08.053653Z digest=sha256:a627938dae275182f2bdcf79482e3690a0e9a50816cd9b09111203e4c4ab7271

Observation 040418a6-dfae-4977-97f0-68ce49ba74f6 · outbound

This paper cites Qsfm: Model pruning based on quantified similarity between feature maps for ai on edge.IEEE Internet of Things Journal, 9(23):24506–24515,.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Qsfm: Model pruning based on quantified similarity between feature maps for ai on edge.IEEE Internet of Things Journal, 9(23):24506–24515,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:10.741796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:08.136865Z digest=sha256:03fb10c536e3718e243f75882e054eaa8e8a9d981fcf135d372a39aa9c7e64d9

Observation 63559539-8457-4376-a2e8-b9030fb3dafa · outbound

This paper cites Joint token pruning and squeezing towards more aggressive compression of vision transformers.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Joint token pruning and squeezing towards more aggressive compression of vision transformers

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:10.586451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:08.203119Z digest=sha256:cb785a0e019e5ed1a98ea14e3198af110c2f10464b724ce23cd86031fedbc409

Observation 32e2e7f0-0d15-4941-a2d5-1f0e6db7e66d · outbound

This paper cites PPT: Token Pruning and Pooling for Efficient Vision Transformers.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration PPT: Token Pruning and Pooling for Efficient Vision Transformers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:08.280794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:08.280794Z digest=sha256:1c1a5e0e6e77d8250c3902d31cdd7c8a799be68dfcdef80035ab46ac841c8198

Observation eca5658d-1943-4dc2-8b1a-8f62dab531a8 · outbound

This paper cites Evo-vit: Slow-fast token evolution for dynamic vision transformer.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Evo-vit: Slow-fast token evolution for dynamic vision transformer

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:10.423172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:08.344854Z digest=sha256:d7c08aa32b5bb288f998559fcd413eb02a5670888066d8fa7e78f8762502b49c

Observation 72664a1a-4b25-4c08-bda6-24766bf60a3e · outbound

This paper cites Global vision transformer pruning with hessian-aware saliency.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Global vision transformer pruning with hessian-aware saliency

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:10.254561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:08.423711Z digest=sha256:dfe3ae7a6db94befd0135c71f97da6d5ea1cc5e60c60572f7fff75220e0082b3

Observation 7c47a347-b9f8-4c80-99c0-4ebe1b3fd104 · outbound

This paper cites A-vit: Adaptive tokens for efficient vision transformer.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration A-vit: Adaptive tokens for efficient vision transformer

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:10.129384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:08.504518Z digest=sha256:c7fcd7533de962acea87b612d9e46c6dc15fab45f1a5f5f65e394f74f8192483

Observation 63afa598-10e5-4f57-88f8-492bde008644 · outbound

This paper cites Tokens-to-token vit: Training vision transformers from scratch on imagenet.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Tokens-to-token vit: Training vision transformers from scratch on imagenet

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:10.005623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:08.575250Z digest=sha256:9ff3ca9e4f7c5480ebf947bf5e102922b8bcb6cb4daa0618f790d3424fc01f22

Observation 2e68c206-74c2-4a0f-8819-3c2255aeadc6 · outbound

This paper cites Vision trans- former with progressive sampling.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Vision trans- former with progressive sampling

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:08.661650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:08.661650Z digest=sha256:37dd1c8824cda5fd5a0572aa44e32ab09379645a46b01afae19f2428125c327e

Observation a6d25efe-71d6-421f-a3f2-91d162418698 · outbound

This paper cites M2m-tag: Training-free many- to-many token aggregation for vision transformer accelera- tion.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration M2m-tag: Training-free many- to-many token aggregation for vision transformer accelera- tion

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:09.878326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:08.730983Z digest=sha256:087ee50aa8ce8faca0214a6a19940c2ea058346afe440a990cef75e8f28dc228

Observation 26b79f09-84bc-4d47-93d4-63977eb608fd · outbound

This paper cites Parameter efficient merging for multimodal large language models with complementary parameter adaptation.arXiv preprint arXiv:2502.17159, 2025.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Parameter efficient merging for multimodal large language models with complementary parameter adaptation.arXiv preprint arXiv:2502.17159, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:08.831348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:08.831348Z digest=sha256:e0a62e686485e27b7c6d476833f3b9a1c913fad9198022ffa7305ad21b10731a

Observation 826c9cc8-af54-414b-8c2e-909cd64c8e08 · outbound

This paper cites DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:09.058019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:09.058019Z digest=sha256:81c635279d403911054e98def4cb54541e0803776e82aafbf06afca7d144c081

Observation 9079dc22-0f25-4fa9-87b7-94dd8acfa084 · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.International Journal of Computer Vision, 127:302–321, 2019.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Semantic under- standing of scenes through the ade20k dataset.International Journal of Computer Vision, 127:302–321, 2019

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:09.750518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T10:29:09.281677Z digest=sha256:8bd17f2d916174c3f4916c4f194df64d1cc2ef3ac4a707d4e0582338f828eec6

Pith citing papers

Observation cd619d8b-e160-4609-ba92-0965376c5715 · inbound

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond cites this paper.

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration

Reference 243

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:36.389460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:36.389460Z digest=sha256:0e896e7961fc8f6693b11daefa2c3aa70b97bbdda23bc509fa7f43cff7c8cfd2

Observation 176cd46b-7d57-4bd3-b913-244b12844db4 · inbound

MaMe & MaRe: Matrix-Based Token Merging and Restoration for Efficient Visual Perception and Synthesis cites this paper.

MaMe & MaRe: Matrix-Based Token Merging and Restoration for Efficient Visual Perception and Synthesis Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:26.675894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T13:06:09.392876Z digest=sha256:41c1cffd1c0401437c064f5c6866ff95cb6df2f556d5988f48182d0051a6e0b3

Observation 3d6a8040-3b49-477c-ba34-4d78edc3078a · inbound

Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors cites this paper.

Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:35:18.966902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T11:31:52.126027Z digest=sha256:0e4c8669c1437c86b385d44e3130734e5a9c03e0ed44581affd8227ff549fef4