Pith. sign in

Paper Citation Record · LEDGER

How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2106.10270.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.10270 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:03:50.087516Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T18:15:59.170440Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c20ce74c-047f-410c-ac72-52b7bcc383ed · inbound

Sigmoid Loss for Language Image Pre-Training cites this paper.

Sigmoid Loss for Language Image Pre-Training How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:05:36.491069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T13:05:36.460932Z digest=sha256:f0a0ac1a61fefde949ee262f2506f9cbc8c13099c8ea2f5aaead2d69b504b0a2

Observation 88d9375c-4005-4576-9c64-077b5c5fa2bc · inbound

Demystifying CLIP Data cites this paper.

Demystifying CLIP Data How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:20:20.323505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T09:20:20.143143Z digest=sha256:fc3d67a4ebef8abc1eac5de0e048b1e751d710358cd3a98c6d3cceb329d18355

Observation 1a7378d4-cace-4508-b514-cf3741ec5691 · inbound

Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey cites this paper.

Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 185

Resolution
verified exact
arxiv_id, observed 2026-05-13T11:32:36.926712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T11:32:36.738536Z digest=sha256:1704c6e64b7be57d16067281ff5a1aee94e5be0a7f4d0ac3f1a1d2f5e3ebc286

Observation 0a682dcd-0704-4295-8db7-3ac96b14c373 · inbound

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control cites this paper.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:38:24.532658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:970f3c68fef92da0bf8953b3a5b5d5ee0d205842e2b962fdcf817098c86be9b2

Observation 2ff296f9-cc39-425d-85f6-cb943fced636 · inbound

DiTASK: Multi-Task Fine-Tuning with Diffeomorphic Transformations cites this paper.

DiTASK: Multi-Task Fine-Tuning with Diffeomorphic Transformations How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T17:03:50.087516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:03:50.087516Z digest=sha256:c136f95cb207fc1231184a606e5287f5c062bc1d1928eafbd1bc9553ddf55042

Observation 69e3555e-2820-42fc-8470-3b0b8cea1dab · inbound

TEMSET-24K: Densely Annotated Dataset for Indexing Multipart Endoscopic Videos using Surgical Timeline Segmentation cites this paper.

TEMSET-24K: Densely Annotated Dataset for Indexing Multipart Endoscopic Videos using Surgical Timeline Segmentation How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:59.632356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:37:59.632356Z digest=sha256:413d34b2bc6b2569cf12f58f19f4b09543c10e868aeb59cbaa392c1ee27872f8

Observation 255794bd-9d13-46bb-a48b-df6496b4b83c · inbound

Scaling Pre-training to One Hundred Billion Data for Vision Language Models cites this paper.

Scaling Pre-training to One Hundred Billion Data for Vision Language Models How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T12:12:30.787739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:12:30.787739Z digest=sha256:ab2018acf64a5f8989088f3352e204420bae0cbae5d96ca427d74d719ca7c6ff

Observation 9bc80515-111b-4fcf-8c17-1f8e21331cc1 · inbound

EMRA-proxy: Enhancing Multi-Class Region Semantic Segmentation in Remote Sensing Images with Attention Proxy cites this paper.

EMRA-proxy: Enhancing Multi-Class Region Semantic Segmentation in Remote Sensing Images with Attention Proxy How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:53.729509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:53.729509Z digest=sha256:71feb70aa506d6e958b65e7c7d49e8a646b31d39323a0fc87bd1e4e742b0fa82

Observation 9fc44e78-5762-44f8-b03f-022e60c60f95 · inbound

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs cites this paper.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:47.047930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:47.047930Z digest=sha256:16a713823ff18ca57fdae33011ff27afacf8ab1cc87ff8131fa1c2eaa230801d

Observation 764911ee-d831-4bb5-84c0-49aa58cfe182 · inbound

Leaner Transformers: More Heads, Less Depth cites this paper.

Leaner Transformers: More Heads, Less Depth How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:12.281412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:12.281412Z digest=sha256:6e6b3f2e77cf2f6b354e7dd28464c5f8ee9968b809dd5af0c6efad8987e9c531

Observation 42585f21-2d84-489f-8b2a-e98b3c9bf0e0 · inbound

Moment Alignment: Unifying Gradient and Hessian Matching for Domain Generalization cites this paper.

Moment Alignment: Unifying Gradient and Hessian Matching for Domain Generalization How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:48:35.044872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:48:35.044872Z digest=sha256:2b94889702f0fde3986480fbaec07267839bec4d36cc61aa3032ab6ecb1a5ad4

Observation 96bd96aa-b993-4134-9747-392bb262feb0 · inbound

Hidden in plain sight: VLMs overlook their visual representations cites this paper.

Hidden in plain sight: VLMs overlook their visual representations How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.940685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.940685Z digest=sha256:c94ada90c5040c8099f5d5003589f103052ad2fdf174953befb58dd235b76e60

Observation a7c6cb3a-4163-4dd1-afa1-20b9344c294f · inbound

InceptionMamba: An Efficient Hybrid Network with Large Band Convolution and Bottleneck Mamba cites this paper.

InceptionMamba: An Efficient Hybrid Network with Large Band Convolution and Bottleneck Mamba How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:07:14.444800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:07:14.444800Z digest=sha256:fa2dddee9917a3ad53da4940b093745c329896d4bce9011e802513bde602d867

Observation 3534d411-7f0c-4e82-8175-a40d15f136fb · inbound

ReStNet: A Reusable & Stitchable Network for Dynamic Adaptation on IoT Devices cites this paper.

ReStNet: A Reusable & Stitchable Network for Dynamic Adaptation on IoT Devices How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:44:09.801949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:44:09.801949Z digest=sha256:2395681c5d96b016bdf3a0f085a3295ef1c66af9fedd69f3ba3cd5f5310df795

Observation f3e75e37-b0af-4a03-8caf-d19a2825939c · inbound

DeepTraverse: A Depth-First Search Inspired Network for Algorithmic Visual Understanding cites this paper.

DeepTraverse: A Depth-First Search Inspired Network for Algorithmic Visual Understanding How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:06.578252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:06.578252Z digest=sha256:495fc1d34434f6a7988ec5bc55dc3ab092ff439a3b5e8febf29e24ede12919f4

Observation 7c5922de-eae6-47b2-9b43-a83d6eb5ce41 · inbound

Pose Matters: Evaluating Vision Transformers and CNNs for Human Action Recognition on Small COCO Subsets cites this paper.

Pose Matters: Evaluating Vision Transformers and CNNs for Human Action Recognition on Small COCO Subsets How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:07.020403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:08:07.020403Z digest=sha256:dc001497209967867a1a8e163ad5420dba3c21cbbfb3f8b4338444919f3740cd

Observation 6b183a73-cc2c-4838-b1c5-c9976c695e2e · inbound

Underwater Monocular Metric Depth Estimation: Real-World Benchmarks and Synthetic Fine-Tuning with Vision Foundation Models cites this paper.

Underwater Monocular Metric Depth Estimation: Real-World Benchmarks and Synthetic Fine-Tuning with Vision Foundation Models How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:40:38.631360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:40:38.631360Z digest=sha256:990bb9f2526d6374f36fc594e12717f5200a67a8a736e7f9ab361a78d96072e6

Observation a2560c90-33d3-4495-af18-358e1fb372a8 · inbound

Elastic ViTs from Pretrained Models without Retraining cites this paper.

Elastic ViTs from Pretrained Models without Retraining How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T09:03:09.342772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:03:09.342772Z digest=sha256:70f43eb3dd9b8b59c502af22b9c40ee09626e3120f6fb7e8739e43909d3e80ec

Observation 272e2d83-974f-48ea-9120-5bf5c4981d82 · inbound

CoMViT: An Efficient Vision Backbone for Supervised Classification in Medical Imaging cites this paper.

CoMViT: An Efficient Vision Backbone for Supervised Classification in Medical Imaging How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T07:01:12.434996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:01:12.434996Z digest=sha256:e9173bd40329ad102c98d0893d0077e0f0cf79dad61b9c82f0aaed144495b729

Observation f6043aef-f7c3-442b-9bf2-838d20253c18 · inbound

Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment cites this paper.

Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:16.095023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:34:16.095023Z digest=sha256:a912714c5e83e02822a2b3e15607c8cddc08ff4390cd1c7a98638b5871dcc2b0

Observation f42dac14-2c5d-4b95-924e-1dc28f5ab6fc · inbound

Causal Attribution via Activation Patching cites this paper.

Causal Attribution via Activation Patching How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:10:02.082743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T11:09:50.326389Z digest=sha256:53203cfcaf7e2b0399c5649f72c0e4978e571f7a5dab3551b67f9021f253410b

Observation 376ee791-ad4f-4850-a5d1-1d93161eae35 · inbound

Human-like Object Grouping in Self-supervised Vision Transformers cites this paper.

Human-like Object Grouping in Self-supervised Vision Transformers How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T21:34:48.709465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T21:34:48.709465Z digest=sha256:a3deb89f1fc3e5eb7b75aef70c07c8d76ff6ee7555be561c1ff126afb23f0f5b

Observation a1b08e9a-7280-4fcb-9bf2-3e635174cd1d · inbound

Decision-Aware Attention Propagation for Vision Transformer Explainability cites this paper.

Decision-Aware Attention Propagation for Vision Transformer Explainability How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:01:01.755658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T04:23:04.096579Z digest=sha256:1eaeba22dea5998f2ccd277288ce587aa4e6a4c66264df8249d49a4297da400d

Observation b3c06afe-72a1-4860-af0a-567a789be8ec · inbound

Enjoy Your Layer Normalization with the Computational Efficiency of RMSNorm cites this paper.

Enjoy Your Layer Normalization with the Computational Efficiency of RMSNorm How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:33:32.625105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T02:29:49.803834Z digest=sha256:9c3b716e71e91213f49faff01a70de8afc5a8349b05b016eff91bc6f8315a504

Observation 7494ba04-06cb-448f-aa2a-656cd19c5252 · inbound

ASAP: Attention Sink Anchored Pruning cites this paper.

ASAP: Attention Sink Anchored Pruning How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:14:45.564776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T08:12:13.451406Z digest=sha256:1e60597992cc15823f43c3d2c654369a477a14c102f896a6aea1eb259282845b

Observation ae73bce2-3402-4ca3-b4fe-61ef2d1c7484 · inbound

Weierstrass Positional Encoding for Vision Transformers cites this paper.

Weierstrass Positional Encoding for Vision Transformers How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:50:23.743705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:48:36.533090Z digest=sha256:6730b5ae929cf38eb99cdf3cc6a76476c36aa60b299883bb1046cf5dce4fdcdf

Observation 20056288-d1b1-4b65-933b-1b176bd3358e · inbound

Large Language Model Teaches Visual Students: Cross-Modality Transfer of Fine-Grained Conceptual Knowledge cites this paper.

Large Language Model Teaches Visual Students: Cross-Modality Transfer of Fine-Grained Conceptual Knowledge How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:15:59.172041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T02:07:27.600563Z digest=sha256:3bb09dca2c8b7a5b879f1c8f242143171dbf37692ed0bbe50467df2176a4d1a9

Observation ebba460c-3cad-45ec-bfdd-c603b8daea9e · inbound

Screening Is Effective for Visual Recognition cites this paper.

Screening Is Effective for Visual Recognition How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T03:08:45.376016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:08:45.376016Z digest=sha256:d1edc62f0a00279c3b4cb358f6a2e155aac5f58076f10927036f0d5683caa0e7

Observation cbcb636e-fb69-4db7-89e7-e96e75eb7dac · inbound

Advancing Multimodal Fusion on Heterogeneous Medical Data with Hybrid Geometry Attention cites this paper.

Advancing Multimodal Fusion on Heterogeneous Medical Data with Hybrid Geometry Attention How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T13:33:15.789562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:33:15.789562Z digest=sha256:0baa7f57da77e3688fec22c4dda74ab545afba20a7507afa0382260a9be621ef

Observation f1edc681-1762-4852-a8af-99802565b854 · inbound

Color Fundus Photography Analysis: Co-evolution of Data, Preprocessing, and Modeling toward Multimodal AI cites this paper.

Color Fundus Photography Analysis: Co-evolution of Data, Preprocessing, and Modeling toward Multimodal AI How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 186

Resolution
unresolved
no resolver link, observed 2026-07-31T23:27:36.860612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:27:36.860612Z digest=sha256:4058d8b95c7f0a5b6bf06fed69e40f50bd33813c4af49401f8cb7a5a3452c3d2

Observation 8a185ae9-54d4-4b74-bfab-00f1a50c0bd0 · inbound

Representation Trajectories Matters: Complementary Evidence for OOD Detection and Image Classification cites this paper.

Representation Trajectories Matters: Complementary Evidence for OOD Detection and Image Classification How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T13:20:38.168263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T13:20:38.168263Z digest=sha256:c25588695ea29f46306cac725d13be6c29afc20eae3d4a71f6e0f3fc87cf0485