{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2022:OQSOZJKMNLTID2TP7ZMQTZ4LB7","short_pith_number":"pith:OQSOZJKM","schema_version":"1.0","canonical_sha256":"7424eca54c6ae681ea6ffe5909e78b0fe41eb9ff5e1bbd602b6163b8861d0429","source":{"kind":"arxiv","id":"2202.07800","version":2},"attestation_state":"computed","paper":{"title":"Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.LG"],"primary_cat":"cs.CV","authors_text":"Chongjian Ge, Jue Wang, Pengtao Xie, Yibing Song, Youwei Liang, Zhan Tong","submitted_at":"2022-02-16T00:19:42Z","abstract_excerpt":"Vision Transformers (ViTs) take all the image patches as tokens and construct multi-head self-attention (MHSA) among them. Complete leverage of these image tokens brings redundant computations since not all the tokens are attentive in MHSA. Examples include that tokens containing semantically meaningless or distractive image backgrounds do not positively contribute to the ViT predictions. In this work, we propose to reorganize image tokens during the feed-forward process of ViT models, which is integrated into ViT during training. For each forward inference, we identify the attentive image tok"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2202.07800","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CV","submitted_at":"2022-02-16T00:19:42Z","cross_cats_sorted":["cs.LG"],"title_canon_sha256":"c2914e2c5997263941d332e0d5846a57e44cfeaea1d4e9c693ae94bd0e0e6053","abstract_canon_sha256":"1f9314b4f984bb2c5e23d4e89f5c997b69aa9c453a296fcd18459a5045267178"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T04:14:41.928986Z","signature_b64":"ndHKu96t5krqsq2JZ/jrPAlXodJz+v+4KGS5fyWjIcMcuJg2UmnCZugzpTAZLy01QGChPcVGVcccpJfYZVKNAg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"7424eca54c6ae681ea6ffe5909e78b0fe41eb9ff5e1bbd602b6163b8861d0429","last_reissued_at":"2026-07-05T04:14:41.928483Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T04:14:41.928483Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.LG"],"primary_cat":"cs.CV","authors_text":"Chongjian Ge, Jue Wang, Pengtao Xie, Yibing Song, Youwei Liang, Zhan Tong","submitted_at":"2022-02-16T00:19:42Z","abstract_excerpt":"Vision Transformers (ViTs) take all the image patches as tokens and construct multi-head self-attention (MHSA) among them. Complete leverage of these image tokens brings redundant computations since not all the tokens are attentive in MHSA. Examples include that tokens containing semantically meaningless or distractive image backgrounds do not positively contribute to the ViT predictions. In this work, we propose to reorganize image tokens during the feed-forward process of ViT models, which is integrated into ViT during training. For each forward inference, we identify the attentive image tok"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2202.07800","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2202.07800/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2202.07800","created_at":"2026-07-05T04:14:41.928558+00:00"},{"alias_kind":"arxiv_version","alias_value":"2202.07800v2","created_at":"2026-07-05T04:14:41.928558+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2202.07800","created_at":"2026-07-05T04:14:41.928558+00:00"},{"alias_kind":"pith_short_12","alias_value":"OQSOZJKMNLTI","created_at":"2026-07-05T04:14:41.928558+00:00"},{"alias_kind":"pith_short_16","alias_value":"OQSOZJKMNLTID2TP","created_at":"2026-07-05T04:14:41.928558+00:00"},{"alias_kind":"pith_short_8","alias_value":"OQSOZJKM","created_at":"2026-07-05T04:14:41.928558+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":25,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.06565","citing_title":"ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation","ref_index":64,"is_internal_anchor":true},{"citing_arxiv_id":"2606.27161","citing_title":"TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference","ref_index":85,"is_internal_anchor":false},{"citing_arxiv_id":"2606.19932","citing_title":"Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models","ref_index":46,"is_internal_anchor":false},{"citing_arxiv_id":"2606.18439","citing_title":"RegimeVGGT: Layer-Wise Spatially Preserving Redundancy Removal for Visual Geometry Grounded Transformer","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2606.12412","citing_title":"Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models","ref_index":39,"is_internal_anchor":false},{"citing_arxiv_id":"2606.12023","citing_title":"ViT-FREE: Efficient Face Recognition via Early Exiting and Synthetic Adaptation","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2606.27660","citing_title":"MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2606.03569","citing_title":"When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2606.02735","citing_title":"See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs","ref_index":40,"is_internal_anchor":false},{"citing_arxiv_id":"2606.27660","citing_title":"MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2605.31457","citing_title":"VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2605.17837","citing_title":"Temporal Aware Pruning for Efficient Diffusion-based Video Generation","ref_index":56,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22372","citing_title":"ASAP: Attention Sink Anchored Pruning","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2605.17837","citing_title":"Temporal Aware Pruning for Efficient Diffusion-based Video Generation","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2505.23617","citing_title":"One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory","ref_index":35,"is_internal_anchor":false},{"citing_arxiv_id":"2510.18091","citing_title":"Accelerating Vision Transformers with Adaptive Patch Sizes","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14110","citing_title":"SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detection","ref_index":28,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14191","citing_title":"CoReDiT: Spatial Coherence-Guided Token Pruning and Reconstruction for Efficient Diffusion Transformers","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10748","citing_title":"Provable Sparse Inversion and Token Relabel Enhanced One-shot Federated Learning with ViTs","ref_index":41,"is_internal_anchor":false},{"citing_arxiv_id":"2605.05848","citing_title":"VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2605.05848","citing_title":"VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2604.13432","citing_title":"MaMe & MaRe: Matrix-Based Token Merging and Restoration for Efficient Visual Perception and Synthesis","ref_index":4,"is_internal_anchor":false},{"citing_arxiv_id":"2604.14563","citing_title":"Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors","ref_index":25,"is_internal_anchor":false},{"citing_arxiv_id":"2604.16745","citing_title":"Why Training-Free Token Reduction Collapses: The Inherent Instability of Pairwise Scoring Signals","ref_index":32,"is_internal_anchor":false},{"citing_arxiv_id":"2604.16854","citing_title":"Certainty Is Redundant: Token Sparsification for Efficient Camouflaged Object Detection with Vision Foundation Models","ref_index":24,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/OQSOZJKMNLTID2TP7ZMQTZ4LB7","json":"https://pith.science/pith/OQSOZJKMNLTID2TP7ZMQTZ4LB7.json","graph_json":"https://pith.science/api/pith-number/OQSOZJKMNLTID2TP7ZMQTZ4LB7/graph.json","events_json":"https://pith.science/api/pith-number/OQSOZJKMNLTID2TP7ZMQTZ4LB7/events.json","paper":"https://pith.science/paper/OQSOZJKM"},"agent_actions":{"view_html":"https://pith.science/pith/OQSOZJKMNLTID2TP7ZMQTZ4LB7","download_json":"https://pith.science/pith/OQSOZJKMNLTID2TP7ZMQTZ4LB7.json","view_paper":"https://pith.science/paper/OQSOZJKM","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2202.07800&json=true","fetch_graph":"https://pith.science/api/pith-number/OQSOZJKMNLTID2TP7ZMQTZ4LB7/graph.json","fetch_events":"https://pith.science/api/pith-number/OQSOZJKMNLTID2TP7ZMQTZ4LB7/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/OQSOZJKMNLTID2TP7ZMQTZ4LB7/action/timestamp_anchor","attest_storage":"https://pith.science/pith/OQSOZJKMNLTID2TP7ZMQTZ4LB7/action/storage_attestation","attest_author":"https://pith.science/pith/OQSOZJKMNLTID2TP7ZMQTZ4LB7/action/author_attestation","sign_citation":"https://pith.science/pith/OQSOZJKMNLTID2TP7ZMQTZ4LB7/action/citation_signature","submit_replication":"https://pith.science/pith/OQSOZJKMNLTID2TP7ZMQTZ4LB7/action/replication_record"}},"created_at":"2026-07-05T04:14:41.928558+00:00","updated_at":"2026-07-05T04:14:41.928558+00:00"}