Pith. sign in

Paper Citation Record · LEDGER

What do Vision Transformers Learn? A Visual Exploration

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2212.06727.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.06727 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:40:32.294790Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:17:14.971022Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c538e3c6-04d4-4d8a-9823-7cd63c719167 · inbound

ResidualDroppath: Enhancing Feature Reuse over Residual Connections cites this paper.

ResidualDroppath: Enhancing Feature Reuse over Residual Connections What do Vision Transformers Learn? A Visual Exploration

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T20:40:32.294790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:40:32.294790Z digest=sha256:478455429c8a3ea14d9461b69c6cc9fe61ee5447e8611cb30be32d291693d20d

Observation 2cce258a-b3d5-4d64-8f1f-73f2da46dd00 · inbound

TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning cites this paper.

TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning What do Vision Transformers Learn? A Visual Exploration

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:15.705533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:11:15.705533Z digest=sha256:546d06b7097893a17974c3af096207be7c225d563b8540a5471e0b77b629335b

Observation 4d7426b4-84cc-4554-8330-4ebe450de04f · inbound

Causal Graphical Models for Vision-Language Compositional Understanding cites this paper.

Causal Graphical Models for Vision-Language Compositional Understanding What do Vision Transformers Learn? A Visual Exploration

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.328119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.328119Z digest=sha256:ed6b96fdc3d7937e196251f58fc26206c9cc1fc6b314034778838aadc24691b4

Observation f05b5391-7e4d-468f-8a6c-73c5fd4584f4 · inbound

Memory Efficient Matting with Adaptive Token Routing cites this paper.

Memory Efficient Matting with Adaptive Token Routing What do Vision Transformers Learn? A Visual Exploration

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:47:56.053994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:47:56.053994Z digest=sha256:b7f331f1dd67170d51db2757359b78dd592ab089dc7b6fb401b8c210fc19aaf7

Observation 4e5bdf5b-22c8-4752-8bf7-ef7ad3bce46a · inbound

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation cites this paper.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation What do Vision Transformers Learn? A Visual Exploration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.581153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.581153Z digest=sha256:494ddf13195c5633bfb9f3229d75174a2dfb0d04648027cf680df55d06738b73

Observation 4a94a302-29d5-4c08-a23d-c8a26386d7f4 · inbound

Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights cites this paper.

Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights What do Vision Transformers Learn? A Visual Exploration

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T21:02:01.001570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:02:01.001570Z digest=sha256:e4efa3c1bbe6fd98834a7d6007fbe708ced56b983e9558db130860dc06231cfd

Observation af3d2fe9-75c3-4303-b35a-1ba8b32c0ed8 · inbound

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models cites this paper.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models What do Vision Transformers Learn? A Visual Exploration

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.604755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.604755Z digest=sha256:42cd2bcd56d001994e5818260c967c7177f2dbc8fa2fac02646f368d260f0b7c

Observation 87bc3f2b-8620-4f94-b1d9-726b1b30fdc6 · inbound

Explainability for Vision Foundation Models: A Survey cites this paper.

Explainability for Vision Foundation Models: A Survey What do Vision Transformers Learn? A Visual Exploration

Reference 253

Resolution
unresolved
no resolver link, observed 2026-08-10T17:26:35.983708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:26:35.983708Z digest=sha256:e1c0cf90454f9b062c188bf353d996378b402c75931f91ab62c2c1fa3adf73e1

Observation 71ebff2d-17a0-4d8a-b8e2-2956bc460053 · inbound

ToSA: Token Merging with Spatial Awareness cites this paper.

ToSA: Token Merging with Spatial Awareness What do Vision Transformers Learn? A Visual Exploration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.883551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.883551Z digest=sha256:cc33a03192ff73add0d82dea7188ae3a8a45e8778426014469960a2ef8d91df4

Observation 98349d3c-4cf6-4b1e-bfa9-44ccd5c843da · inbound

Towards Distributed Neural Architectures cites this paper.

Towards Distributed Neural Architectures What do Vision Transformers Learn? A Visual Exploration

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.008165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.008165Z digest=sha256:d86941fd5d3e30ddbcdd2e7635e758efb8794cca60466309dd2f5cad5bcecfee

Observation 86964e6e-d519-4f8d-a782-f5c8373628f6 · inbound

Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptation cites this paper.

Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptation What do Vision Transformers Learn? A Visual Exploration

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:05.620893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:05.620893Z digest=sha256:0e2891d89afef7a19fdd909e09b5781a580511397ab8fa3acbbf96e3053b7eda

Observation adb8bb85-54dd-4a31-bbc4-d2298cb69104 · inbound

A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models cites this paper.

A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models What do Vision Transformers Learn? A Visual Exploration

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:46:14.355326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-12T01:41:44.898358Z digest=sha256:ee43f9ad596cd28ce2c0b7a33c94a9e5ed9e1a93625e50db1d14de534c457964

Observation f15d0c8f-f091-4204-bd86-89d553fcb327 · inbound

Realiz3D: 3D Generation Made Photorealistic via Domain-Aware Learning cites this paper.

Realiz3D: 3D Generation Made Photorealistic via Domain-Aware Learning What do Vision Transformers Learn? A Visual Exploration

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:39:50.416786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T07:36:16.124800Z digest=sha256:7117e8dfb773a453466e04fcd7b74b2f82cae2413f39d1e18fac2f80c2631abb

Observation 0fe62a99-234c-4b4f-a25d-e6c56e45e090 · inbound

LESSViT: Robust Hyperspectral Representation Learning under Spectral Configuration Shift cites this paper.

LESSViT: Robust Hyperspectral Representation Learning under Spectral Configuration Shift What do Vision Transformers Learn? A Visual Exploration

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:18:13.744646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T11:16:01.428052Z digest=sha256:37b213c9ae54e51fa06fc7aba593a06aa9f544e33dc469033deccd9dc1adca16

Observation 54c4e797-6c7b-49d6-a5d0-fce4f59bd905 · inbound

Textual Supervision Enhances Geospatial Representations in Vision-Language Models cites this paper.

Textual Supervision Enhances Geospatial Representations in Vision-Language Models What do Vision Transformers Learn? A Visual Exploration

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:17:14.972599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T22:06:12.034163Z digest=sha256:c0758f7c021dafbfe9e803e553c733418eb1f96fead7baa82e7fb8953b211b1f

Observation c3b693d2-8507-44b3-b09f-9dcea39679c4 · inbound

Token-Based Affordance Grounding with Large Vision-Language Models cites this paper.

Token-Based Affordance Grounding with Large Vision-Language Models What do Vision Transformers Learn? A Visual Exploration

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T01:18:51.054590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:18:51.054590Z digest=sha256:a030668fd332b453c05221c26544510601c02bea9d15b75d85138a69fcdd053d