Pith. sign in

Paper Citation Record · LEDGER

Visual Transformers: Token-based Image Representation and Processing for Computer Vision

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2006.03677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2006.03677 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:37:14.547591Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:09:57.476599Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cccd1c63-7399-4300-8963-5eab62fb8407 · inbound

Clustered Patch Embeddings for Permutation-Invariant Classification of Whole Slide Images cites this paper.

Clustered Patch Embeddings for Permutation-Invariant Classification of Whole Slide Images Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T21:37:14.547591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:37:14.547591Z digest=sha256:12fef3e8e760dc60cc78a9942bd591e9136a0edeb71a1a80cf0395be78c06268

Observation 8a59969b-d227-4ad9-819f-4bd529fe7755 · inbound

Recall and Refine: A Simple but Effective Source-free Open-set Domain Adaptation Framework cites this paper.

Recall and Refine: A Simple but Effective Source-free Open-set Domain Adaptation Framework Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T17:28:06.493438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:28:06.493438Z digest=sha256:11616d53b5376a4d53f2eacd7704ebfa7838f444023221f5dafaa1053464ea7e

Observation b21399bc-3865-4602-bfea-d21591e0cf01 · inbound

Exploring Foundation Models Fine-Tuning for Cytology Classification cites this paper.

Exploring Foundation Models Fine-Tuning for Cytology Classification Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T14:43:52.602676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:43:52.602676Z digest=sha256:58b0fb957d6d6026e446f79b3ccb8c705356c828430a349f1a7e9afb10f83d62

Observation 10c16e01-2d67-488d-89b5-6743e927b96f · inbound

Real-Time Anomaly Detection in Video Streams cites this paper.

Real-Time Anomaly Detection in Video Streams Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-12T05:59:46.284846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:59:46.284846Z digest=sha256:df388838577f95c35b77d8f5d2b3d8c7a1ee17e8a902e0e97f6fea374345ce38

Observation f0b79e23-c841-4b99-ab04-62b89f3e730d · inbound

TSUBF-Net: Trans-Spatial UNet-like Network with Bi-direction Fusion for Segmentation of Adenoid Hypertrophy in CT cites this paper.

TSUBF-Net: Trans-Spatial UNet-like Network with Bi-direction Fusion for Segmentation of Adenoid Hypertrophy in CT Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T05:03:20.700980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:03:20.700980Z digest=sha256:04b00c1c9c4c8f3354a8e6ff572b39fa037e0c97d65ea2745245fb532d578585

Observation 2fd49c31-be75-4af4-84ec-3e9a8d5c17d2 · inbound

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction cites this paper.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.887844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.887844Z digest=sha256:9257140cf950e2619211dda3c5cd1e2f9c4f8a4db988283ffd5a240bf10816f1

Observation 6d7442f7-77be-4bfe-8f4b-23ddc09882af · inbound

Dual Mutual Learning Network with Global-local Awareness for RGB-D Salient Object Detection cites this paper.

Dual Mutual Learning Network with Global-local Awareness for RGB-D Salient Object Detection Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:56.320519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:56.320519Z digest=sha256:41ea79441d033b62afd0ccd566c0aecaafefcb953d1f10520a23febdf8371abf

Observation f85bf048-364a-4f71-931e-2948d57ccca3 · inbound

PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides cites this paper.

PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:47:26.640556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:47:26.640556Z digest=sha256:85a838babf5937a35d8e16d8ed11c5716211908baad57ced962e2f61b1fb361a

Observation 36094c41-ccbc-43f4-8e2f-5ecdca2e55ac · inbound

BRIGHT-VO: Brightness-Guided Hybrid Transformer for Visual Odometry with Multi-modality Refinement Module cites this paper.

BRIGHT-VO: Brightness-Guided Hybrid Transformer for Visual Odometry with Multi-modality Refinement Module Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:25:59.538557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:25:59.538557Z digest=sha256:525d56334d72b87838332d4dc9b8590a1d21ff0413b91b0952c9c6b3ae267a1d

Observation abb19d1f-e528-4406-97a0-69ecaa6e1b87 · inbound

Zero-Shot Warning Generation for Misinformative Multimodal Content cites this paper.

Zero-Shot Warning Generation for Misinformative Multimodal Content Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T17:53:24.593745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:53:24.593745Z digest=sha256:af61581683390aec187a1166de6284d2ccc8ad7fd59893d46e44a65e8c7c4765

Observation c9bc989c-6939-4250-b889-573ef6789267 · inbound

Exploring Visual Embedding Spaces Induced by Vision Transformers for Online Auto Parts Marketplaces cites this paper.

Exploring Visual Embedding Spaces Induced by Vision Transformers for Online Auto Parts Marketplaces Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T18:08:52.485744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:08:52.485744Z digest=sha256:13f8bd1068fff7d4f4b08b08c1091eff4187105005b66152c6381a6d2b684cc0

Observation 6577d602-f489-4662-b633-ea4472b74b5f · inbound

Defending against Backdoor Attacks via Module Switching cites this paper.

Defending against Backdoor Attacks via Module Switching Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:32:04.763637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T20:28:24.169272Z digest=sha256:7bf5d997dfcc58ea9dc3e753d6ae5d65566fa0ecca424234551e2f2061fb5c02

Observation e34d4855-55ad-4fd9-b450-aebd848e2419 · inbound

Can Vision Transformers with ResNet's Global Features Fairly Authenticate Demographic Faces? cites this paper.

Can Vision Transformers with ResNet's Global Features Fairly Authenticate Demographic Faces? Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:22:59.772724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:22:59.772724Z digest=sha256:97673992eea58abcc02ca4c82a569a6e6b25493a1c3fbebad282de9fc33aa5ca

Observation b99e02f5-99a0-4e77-a5c0-064d61fb9abf · inbound

Hierarchical Deep Feature Fusion and Ensemble Learning for Enhanced Brain Tumor MRI Classification cites this paper.

Hierarchical Deep Feature Fusion and Ensemble Learning for Enhanced Brain Tumor MRI Classification Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:55:35.034653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:55:35.034653Z digest=sha256:35ad6cdc2dff652a6a0c4b34e10d4223974505e6567f355c445069c5fd290e61

Observation 4c5aeb60-c08f-4187-af7a-0ab4da22106d · inbound

Hybrid Ensemble Approaches: Optimal Deep Feature Fusion and Hyperparameter-Tuned Classifier Ensembling for Enhanced Brain Tumor Classification cites this paper.

Hybrid Ensemble Approaches: Optimal Deep Feature Fusion and Hyperparameter-Tuned Classifier Ensembling for Enhanced Brain Tumor Classification Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:33.510651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:33.510651Z digest=sha256:992263d86cdf156e0aed4d1754cf9d3a36f7efcd93ebd4c501b73bf8be5acaef

Observation 384dad4f-8e88-4557-ab0b-b3e25bf0b2e5 · inbound

On the Performance of Concept Probing: The Influence of the Data (Extended Version) cites this paper.

On the Performance of Concept Probing: The Influence of the Data (Extended Version) Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:21.395406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:21.395406Z digest=sha256:1ccb3c0f5636765c875b8f9187cac8964b1532f2e2099d2d22599cb08bf719bb

Observation e306f25c-a077-4e99-91eb-7661ce9e69f1 · inbound

Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems cites this paper.

Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T12:39:41.651354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:39:41.651354Z digest=sha256:bbaab471830e5891305b82e83d5a43a5d75b3e0285e6252a6edda83baf55de1b

Observation 3a97b8a1-aea6-48be-ae4e-aec3f4f4c7dd · inbound

Pruning and Malicious Injection: A Retraining-Free Backdoor Attack on Transformer Models cites this paper.

Pruning and Malicious Injection: A Retraining-Free Backdoor Attack on Transformer Models Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:42:26.151279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:42:26.151279Z digest=sha256:a42b57e1209e0bd683db89c6867f56ffd447e940601e58375755d36aef46072d

Observation babc0f36-4720-4574-bf17-7d6a1ba7dd3b · inbound

On the Complexity-Faithfulness Trade-off of Gradient-Based Explanations cites this paper.

On the Complexity-Faithfulness Trade-off of Gradient-Based Explanations Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:43.996459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:29:43.996459Z digest=sha256:903c6c2101f72809ee851f198f72a44fa7ace7939b431a5711827a23d3d42b36

Observation 838fa267-a067-487e-889a-15b41a2cd9a2 · inbound

Focus Through Motion: RGB-Event Collaborative Token Sparsification for Efficient Object Detection cites this paper.

Focus Through Motion: RGB-Event Collaborative Token Sparsification for Efficient Object Detection Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T10:40:00.899017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:40:00.899017Z digest=sha256:be943e11a5718956c3e8c4c462d3ca8a3c642c2d0c25cd14d24fe86349022666

Observation 3b8d841e-b0ad-4195-8b38-0b4617977183 · inbound

WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World cites this paper.

WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-03T17:02:40.945788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:02:40.945788Z digest=sha256:2e56ac04356a430f4e9f0b90c8d3648ec04a9393232476ad9a7caeb78cdeaefd

Observation a3c569f9-f741-41f5-ad29-97d634cdb446 · inbound

MoDora: Tree-Based Semi-Structured Document Analysis System cites this paper.

MoDora: Tree-Based Semi-Structured Document Analysis System Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:10:15.528788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T19:09:41.482713Z digest=sha256:2f1dc55dbd3fdd1681d73e57ddde9ecce09be8754e4f083c50d815624b9b65a7

Observation 3c2b3914-1d60-4f3e-98bc-aa448f7b5a53 · inbound

Rethinking Intrinsic Dimension Estimation in Neural Representations cites this paper.

Rethinking Intrinsic Dimension Estimation in Neural Representations Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:54:48.758900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T00:50:49.134622Z digest=sha256:12b0f26c49105987e60fc25e3af23fac5681a5163f466b9fa76319783d9a9bab

Observation 5dfddba1-9fa1-4cb7-9865-0301af09c7ba · inbound

Modeling Subjective Urban Perception with Human Gaze cites this paper.

Modeling Subjective Urban Perception with Human Gaze Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:47:15.929843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-09T19:12:54.038217Z digest=sha256:5f21ce0e1b184e6cb609f7f8726fce2e58c75c4a961a720742bb9cd66b945f18

Observation 3ae6afd6-d479-48c3-9d89-a954915cdb15 · inbound

SAIL: Structure-Aware Interpretable Learning for Anatomy-Aligned Post-hoc Explanations in OCT cites this paper.

SAIL: Structure-Aware Interpretable Learning for Anatomy-Aligned Post-hoc Explanations in OCT Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:41.076077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T18:16:35.927272Z digest=sha256:733f8a4af99b5b1ccc34c4daa45b8f8b17fea1271614cc2bce248fc7f65f0ac1

Observation fb19430e-446d-4162-a391-142c9a35af9e · inbound

M3Net: A Macro-to-Meso-to-Micro Clinical-inspired Hierarchical 3D Network for Pulmonary Nodule Classification cites this paper.

M3Net: A Macro-to-Meso-to-Micro Clinical-inspired Hierarchical 3D Network for Pulmonary Nodule Classification Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:59:27.369856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T20:55:57.769186Z digest=sha256:92a65ff46379979a62040a103d9d3251af4c8dded4389676d727e86083540e0f

Observation aac02284-2aa9-466d-8eb3-fea5bf884b38 · inbound

Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI cites this paper.

Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:43:15.188039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T11:40:19.027987Z digest=sha256:1a8073bce9a1a0c3740b631aa991955f805d1fa9eaf28eb21b3d711629c39107

Observation 608cde3a-c82e-4b76-bedb-069ef682970b · inbound

Deep Attention Reweighting: Post-Hoc Attention-Based Feature Aggregation in CNNs for Disentangling Core and Spurious Features under Spurious Correlations cites this paper.

Deep Attention Reweighting: Post-Hoc Attention-Based Feature Aggregation in CNNs for Disentangling Core and Spurious Features under Spurious Correlations Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:09:38.768756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T05:04:59.177437Z digest=sha256:ea463aa65f96eb037684c823265fba2bf8c70c82eeaf7355ec6b36592285eabd

Observation 098a2994-ae5d-49be-9b8e-8abbe2e85b91 · inbound

Forged Calamity: Benchmark for Cross-Domain Synthetic Disaster Detection in the Age of Diffusion cites this paper.

Forged Calamity: Benchmark for Cross-Domain Synthetic Disaster Detection in the Age of Diffusion Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:07.327251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T21:32:27.296146Z digest=sha256:f8f62d261b395a1c0c120eb3004b53e0fa12d20ee5f87cf5d7bad4bc92844243

Observation e82670d2-c508-4a32-b824-a848a5d8f388 · inbound

Co-occurring associated retained concepts in Diffusion Unlearning cites this paper.

Co-occurring associated retained concepts in Diffusion Unlearning Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:09:57.478300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-26T00:50:58.689718Z digest=sha256:ec9517a540bff016f1671fdc47f67444fea974d29ee93944ec49b9b6f5d57cba

Observation 32837a37-b408-4ae6-9442-6e54b233b313 · inbound

One Framework for All: Cross-Modal Membership Inference for Generative Models cites this paper.

One Framework for All: Cross-Modal Membership Inference for Generative Models Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-11T19:58:05.484186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:58:05.484186Z digest=sha256:c536d1f2e1497956f028a7a7ebfd54f2efdc16edd1ed3fd9bf0f1498a75deeba

Observation cf25f6a4-819f-4763-a036-9e8d143a63d5 · inbound

Towards Effective Federated Multimodal Graph Learning via Navigating Multifaceted Heterogeneity cites this paper.

Towards Effective Federated Multimodal Graph Learning via Navigating Multifaceted Heterogeneity Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T00:38:04.305629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:38:04.305629Z digest=sha256:27330072ed223436a19957a68c93b57b068e23ae7bac15e45452d19a5150a33b