Pith. sign in

Paper Citation Record · LEDGER

An Empirical Study of Training Self-Supervised Vision Transformers

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2104.02057.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.02057 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:40:17.417495Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

17
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f2686462-a443-4b7e-a018-4b22fafdeaaf · inbound

BEiT: BERT Pre-Training of Image Transformers cites this paper.

BEiT: BERT Pre-Training of Image Transformers An Empirical Study of Training Self-Supervised Vision Transformers

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T11:50:11.503940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T11:50:11.476015Z digest=sha256:985122099a84605ce14ce75fc00cda2e4078876b868d1389ba2c2d2480ce61f1

Observation 0db076aa-81db-4f10-995a-66bf1e7b52c6 · inbound

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution cites this paper.

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution An Empirical Study of Training Self-Supervised Vision Transformers

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:12:31.385104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T08:12:30.984870Z digest=sha256:14fe771793fbc8651cb9a9fa09ae90d86ca025572c3f5ff69497c8e27b5da5ae

Observation 503156a9-01f2-46d4-b08e-8768895902d1 · inbound

Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels cites this paper.

Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels An Empirical Study of Training Self-Supervised Vision Transformers

Reference 129

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T16:35:48.149598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T16:35:47.826165Z digest=sha256:9f87f2a9302b78765f8d28befb2b88d76a83b3962c9a6e3e42421b30eb3727b1

Observation 24e23a72-8f7c-487c-accf-cf26de8bbac9 · inbound

Revisiting Feature Prediction for Learning Visual Representations from Video cites this paper.

Revisiting Feature Prediction for Learning Visual Representations from Video An Empirical Study of Training Self-Supervised Vision Transformers

Reference 230

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:40:23.862062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T12:40:23.709098Z digest=sha256:982a6368b7cc35666bf3f8fbea101706641aaffd88216a5c2383a86b96a17905

Observation 7ed57719-9bd5-40fd-9bdd-75fbd04748f4 · inbound

GAIR: Location-Aware Self-Supervised Contrastive Pre-Training with Geo-Aligned Implicit Representations cites this paper.

GAIR: Location-Aware Self-Supervised Contrastive Pre-Training with Geo-Aligned Implicit Representations An Empirical Study of Training Self-Supervised Vision Transformers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:42:13.409272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T22:41:00.027266Z digest=sha256:95c92d00ad412d555990dd570dac9a4c494bcdbfa99a64e756a88b57b61e5bb4

Observation 5c89becb-6785-416b-a98b-ef71a599ad6c · inbound

Self Distillation via Iterative Constructive Perturbations cites this paper.

Self Distillation via Iterative Constructive Perturbations An Empirical Study of Training Self-Supervised Vision Transformers

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:17.417495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:40:17.417495Z digest=sha256:769841f1542f5677f4a770e8c2ea836edcd2b4d07eb5e3103a5cc7be70df031e

Observation 694a59dd-5a4d-40fa-8b3a-53c6246972b8 · inbound

Cut out and Replay: A Simple yet Versatile Strategy for Multi-Label Online Continual Learning cites this paper.

Cut out and Replay: A Simple yet Versatile Strategy for Multi-Label Online Continual Learning An Empirical Study of Training Self-Supervised Vision Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.095167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:41.095167Z digest=sha256:aabfc3736281df9510340c1201c91139bb59c9c6e1d2c33fecebb7528c03e5cb

Observation 97ce9b05-cb7c-4620-98da-8e97e7d10f55 · inbound

An Augmentation-Aware Theory for Self-Supervised Contrastive Learning cites this paper.

An Augmentation-Aware Theory for Self-Supervised Contrastive Learning An Empirical Study of Training Self-Supervised Vision Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:29.010114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:22:29.010114Z digest=sha256:46255b5c308cbac42d6185a02072b5035a0e17a995ec475bbe9944a897293b18

Observation 6bf1eaae-e437-4f15-9902-44650ed9bd71 · inbound

Hidden in plain sight: VLMs overlook their visual representations cites this paper.

Hidden in plain sight: VLMs overlook their visual representations An Empirical Study of Training Self-Supervised Vision Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.866030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.866030Z digest=sha256:95cbbbc09c468f7b84a60ea2362a60d599a75c00c1f51c4ee4fd4859dfc5af6d

Observation 1afc5160-e24b-4743-897b-610538d94610 · inbound

Uncertainty Awareness Enables Efficient Labeling for Cancer Subtyping in Digital Pathology cites this paper.

Uncertainty Awareness Enables Efficient Labeling for Cancer Subtyping in Digital Pathology An Empirical Study of Training Self-Supervised Vision Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:12:48.510234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:12:48.510234Z digest=sha256:15b0b389aa6c10b734bbbb7a7c66c33f634683b69580d7b669e0253c065d32c4

Observation 5c07c8f3-2d4b-41c1-a4d8-0b79ea2e1fb4 · inbound

PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training cites this paper.

PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training An Empirical Study of Training Self-Supervised Vision Transformers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:56:53.139981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T22:53:23.271423Z digest=sha256:0314f4a661ce68d40affb7bc345849c0a272672a15cd626e589c56bbbda2bed1

Observation f6ffccf8-3b6b-48f0-854b-9be69f469b85 · inbound

Self-Soupervision: Cooking Model Soups without Labels cites this paper.

Self-Soupervision: Cooking Model Soups without Labels An Empirical Study of Training Self-Supervised Vision Transformers

Reference 274

Resolution
unresolved
no resolver link, observed 2026-08-03T05:18:30.707009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:18:30.707009Z digest=sha256:0e13058fca56e418526e72b0d98e054062696cad207d7786a52ea1873c9b1c86

Observation 6402da3f-8c90-4237-bb2f-22319c61b382 · inbound

Heuristic Style Transfer for Real-Time, Efficient Weather Attribute Detection cites this paper.

Heuristic Style Transfer for Real-Time, Efficient Weather Attribute Detection An Empirical Study of Training Self-Supervised Vision Transformers

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:40:27.076956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:37:59.894503Z digest=sha256:0d7c83ff6eb7a4b8cd17bb4d0a209bf6c80b092bc2e0caf1974f2f80efba76cb

Observation 342f2900-19dd-4e3b-add3-e470fb2883da · inbound

Style-Based Neural Architectures for Real-Time Weather Classification cites this paper.

Style-Based Neural Architectures for Real-Time Weather Classification An Empirical Study of Training Self-Supervised Vision Transformers

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:51:03.488100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T04:32:29.400543Z digest=sha256:1a63b57f3d0deb083e4c88291cc49ee06809083283f955c59e08f0d87751e7a0

Observation 3c301b2c-af92-463a-bedd-24fb94f32b1c · inbound

Text-Conditional JEPA for Learning Semantically Rich Visual Representations cites this paper.

Text-Conditional JEPA for Learning Semantically Rich Visual Representations An Empirical Study of Training Self-Supervised Vision Transformers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:51:30.647800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T17:58:43.650308Z digest=sha256:76cbdcbe173ca6607112e61aedd9b74b17940fe9924915f5d0a36972aa50c729

Observation a2921392-a880-4238-bf1c-6459a960a07f · inbound

Prognostic Value of Lung Ultrasound Biomarkers for Readmission Risk in Congestive Heart Failure: A Pilot Data-Driven Analysis cites this paper.

Prognostic Value of Lung Ultrasound Biomarkers for Readmission Risk in Congestive Heart Failure: A Pilot Data-Driven Analysis An Empirical Study of Training Self-Supervised Vision Transformers

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:43:26.413714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T15:41:31.181025Z digest=sha256:9195ece270255dfb1c4be63e0dbdad9f9ac0e6c369298cbbf410a569e998a6b9

Observation ba36f626-90bb-4bbf-a811-5e001edf2eaa · inbound

Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting cites this paper.

Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting An Empirical Study of Training Self-Supervised Vision Transformers

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-12T08:12:18.373103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:12:18.373103Z digest=sha256:92e4eeb392b2542d4ac19db095811e58ca4e42a4f9df9f59b3ad6346bd581d65