Pith. sign in

Paper Citation Record · LEDGER

Token Pooling in Vision Transformers

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2110.03860.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2110.03860 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:34:21.948112Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:39:29.498864Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1ea5df7b-2946-42a2-a55a-76c068cfcb39 · inbound

Token Merging: Your ViT But Faster cites this paper.

Token Merging: Your ViT But Faster Token Pooling in Vision Transformers

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T20:52:10.491697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T20:52:10.065925Z digest=sha256:957e736d93204b5ebeba6403626e925d84f8ffc73adc3932eac99b85d904ce54

Observation d0bbac1e-e5eb-4a33-8096-08d06c213a4d · inbound

ImagePiece: Content-aware Re-tokenization for Efficient Image Recognition cites this paper.

ImagePiece: Content-aware Re-tokenization for Efficient Image Recognition Token Pooling in Vision Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T10:34:21.948112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:34:21.948112Z digest=sha256:a6f62610fde55bf9e397b3cd7c13e1a6d923e6a1ef65489360af1f25e33f3368

Observation e1ebb4f2-de57-4697-950f-79e0eddddd7c · inbound

Token Pruning for Caching Better: 9 Times Acceleration on Stable Diffusion for Free cites this paper.

Token Pruning for Caching Better: 9 Times Acceleration on Stable Diffusion for Free Token Pooling in Vision Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:58:47.476902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:58:47.476902Z digest=sha256:4be37a6db015c6ccb88272c35f64b680351590610dd104ce038c3add236f35e7

Observation f081b9ce-b6cd-4af0-9820-9d75f8e2dc0c · inbound

Cached Adaptive Token Merging: Dynamic Token Reduction and Redundant Computation Elimination in Diffusion Model cites this paper.

Cached Adaptive Token Merging: Dynamic Token Reduction and Redundant Computation Elimination in Diffusion Model Token Pooling in Vision Transformers

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:27.372754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:27.372754Z digest=sha256:8b596515a40e9b644a5e92ef4fda1fec3e01ed5c090542e6383efa59ffdcf413

Observation 5881e8b5-c92a-453e-9282-cec7857fda4a · inbound

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models cites this paper.

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models Token Pooling in Vision Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:08.722595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:08:08.722595Z digest=sha256:e372e8133beefe9a15a3436d49d9db77c3975f9c3a8bbacdd67c960bb6e4047b

Observation 788a1717-40dd-4ee4-8aa0-f37e6cdbb4a8 · inbound

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation cites this paper.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Token Pooling in Vision Transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.781218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.781218Z digest=sha256:a367e124e163f878ce49196ae32309fd68200e8673e49037e37aab1ffd314b4c

Observation ea941ac2-f62c-44be-937b-e125dc1e5b6e · inbound

D\'ej\`a Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse cites this paper.

D\'ej\`a Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse Token Pooling in Vision Transformers

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:12.625665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:12.625665Z digest=sha256:68e3aae3d307e53c6a7582e837ba41ca93d4cc949e9f5d761d1e8a0280ade1d4

Observation accbf4d4-a0f0-4cef-9e4c-d95024abc0e1 · inbound

Geo-RepNet: Geometry-Aware Representation Learning for Surgical Phase Recognition in Endoscopic Submucosal Dissection cites this paper.

Geo-RepNet: Geometry-Aware Representation Learning for Surgical Phase Recognition in Endoscopic Submucosal Dissection Token Pooling in Vision Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:10.871048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:10.871048Z digest=sha256:61611f82146da1020541737599db563d61c0325cfec10bb8c03c8d50f05be4d4

Observation ba497e75-eca8-426b-9958-56a0f92e0bc5 · inbound

Training-free Token Reduction for Vision Mamba cites this paper.

Training-free Token Reduction for Vision Mamba Token Pooling in Vision Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T16:16:49.332131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:16:49.332131Z digest=sha256:a0031c8284342c745aa778261e44dcb846f09fa895745453736b1d8151d435f4

Observation cd5d1a64-c8a4-45f3-b6f7-43293047e6dd · inbound

Why Training-Free Token Reduction Collapses: The Inherent Instability of Pairwise Scoring Signals cites this paper.

Why Training-Free Token Reduction Collapses: The Inherent Instability of Pairwise Scoring Signals Token Pooling in Vision Transformers

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:02:25.409065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T07:59:10.678150Z digest=sha256:5a01feb6bddcd3305dea64c507fc4dd52b15b9cd04bc5855e999c4556824d19d

Observation e50f609d-fa0e-4dbe-83f0-2c2191315f8a · inbound

Accelerating Vision Foundation Models with Drop-in Depthwise Convolution cites this paper.

Accelerating Vision Foundation Models with Drop-in Depthwise Convolution Token Pooling in Vision Transformers

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:04:41.468875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T07:03:37.388861Z digest=sha256:f6de949c1f3de525c141f915bdbd69aa85292cf3a3fe810e4da0173869028cf0

Observation 6e6994dc-e9df-4e12-8564-d5f1846809a0 · inbound

RAPID: Layer-Wise Redundancy-Aware Pruning and Importance-Driven Token Merging for Efficient ViT cites this paper.

RAPID: Layer-Wise Redundancy-Aware Pruning and Importance-Driven Token Merging for Efficient ViT Token Pooling in Vision Transformers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:27:24.484522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T19:46:23.188930Z digest=sha256:b64c218ec6b411f2fd8a4b29c472000669e409587f9ac78af134c5318c3d121f

Observation 687b7ebe-694a-4e53-b987-c7eb81cf76dd · inbound

Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models cites this paper.

Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models Token Pooling in Vision Transformers

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:39:29.501847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-26T17:53:38.503877Z digest=sha256:23127caba74bdbf34bc6040955e36981bccaaead7a2937dfded776c147279615

Observation 3d28a755-b435-451d-a2aa-e8fb81fd220e · inbound

DTM-Codec: Dynamic Token Masking for VFR Speech Coding with Efficient Boundary Selection cites this paper.

DTM-Codec: Dynamic Token Masking for VFR Speech Coding with Efficient Boundary Selection Token Pooling in Vision Transformers

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T02:14:11.489247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T02:10:08.691873Z digest=sha256:f2cf36af3a8471ae4a7da926b827f4de27d92f6b35331beb027f4c0a3e093658