Pith. sign in

Paper Citation Record · LEDGER

Learning to Merge Tokens in Vision Transformers

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2202.12015.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2202.12015 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:01:23.804525Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T05:09:36.498887Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c65e594d-c154-445b-a67a-def22ddd5241 · inbound

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation cites this paper.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Learning to Merge Tokens in Vision Transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.804525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.804525Z digest=sha256:00fa6a051ec2c58b78081f830e4f7b62fcff72bfac4b963354ee8c8d6b046765

Observation 06250222-426a-428f-9eb6-6a309faf1aba · inbound

Token Compression Meets Compact Vision Transformers: A Survey and Comparative Evaluation for Edge AI cites this paper.

Token Compression Meets Compact Vision Transformers: A Survey and Comparative Evaluation for Edge AI Learning to Merge Tokens in Vision Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:53:26.621820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:53:26.621820Z digest=sha256:6081c982b22b42c550f37c84834ce161b6407457a7da3808092588c128ec766e

Observation 5bff8415-6a1c-46a9-8a94-14605872bba8 · inbound

Training-free Token Reduction for Vision Mamba cites this paper.

Training-free Token Reduction for Vision Mamba Learning to Merge Tokens in Vision Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T16:16:49.380798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:16:49.380798Z digest=sha256:72d9212424dfe65987d93e964543a5827d4c7cc008174bf8519c74c35614ed27

Observation 78e4a7eb-1f02-45fb-8b56-410e0465d0d5 · inbound

FastVGGT: Training-Free Acceleration of Visual Geometry Transformer cites this paper.

FastVGGT: Training-Free Acceleration of Visual Geometry Transformer Learning to Merge Tokens in Vision Transformers

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:36:06.967355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T23:36:06.870759Z digest=sha256:8168e3918ec3ef33e58e3910141318a19c93afbbbeb5e761b436aaeeecb6af01

Observation 396f89f9-983c-400e-95c0-1009c09e924e · inbound

Detecting Regional Spurious Correlations in Vision Transformers via Token Discarding cites this paper.

Detecting Regional Spurious Correlations in Vision Transformers via Token Discarding Learning to Merge Tokens in Vision Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.375051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:37.375051Z digest=sha256:124973e197c232e564ea44102a53e6c78bf31afd50d5a74d2c84dd3295f0bd4e

Observation 64a6bd45-73c4-433e-8137-78834c6d13f6 · inbound

Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors cites this paper.

Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors Learning to Merge Tokens in Vision Transformers

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:35:18.956741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T11:31:52.126027Z digest=sha256:6e035a17ebf068525ab98bd529cb6848348ec02cbd39bc7a1d1dd04b34058a93

Observation 934e1f69-e30d-4e48-bf54-9d239a0a191a · inbound

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference cites this paper.

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference Learning to Merge Tokens in Vision Transformers

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:29.913691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T18:17:53.013043Z digest=sha256:4f232be4a0ee1c56cc91afdf0688bb6daaf4afdb124c06b3a9a725f1be9dc821

Observation 43b2b66b-a933-4f03-ab63-9851d71c198b · inbound

ConsisFormer: Compute-Efficient Transformer for Wireless Foundation Models Based on Channel Consistency cites this paper.

ConsisFormer: Compute-Efficient Transformer for Wireless Foundation Models Based on Channel Consistency Learning to Merge Tokens in Vision Transformers

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:09:36.501400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T16:16:12.898027Z digest=sha256:766943482654c1b099b3733a30bc3685afd40ae058c8fbad5bfe786d5ebe0299

Observation 1943de30-9f1e-4531-82f7-11f2f0aaf79b · inbound

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs cites this paper.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Learning to Merge Tokens in Vision Transformers

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:22.142467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:e20691fa6fc2b02bfa36e203e99e553d16856a2b5234ea325380c5b05e0a0ff1

Observation f15b98f5-94d5-4933-a1a3-6035195e3a36 · inbound

REDI: Corpus Aware Patch Ranking for DINOv3 Token Reduction cites this paper.

REDI: Corpus Aware Patch Ranking for DINOv3 Token Reduction Learning to Merge Tokens in Vision Transformers

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:55:41.643978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T05:57:55.575289Z digest=sha256:10323d636b72336b183a1c232130a87b1e695330703271fce30d2bca4e9e04d9