Pith. sign in

Paper Citation Record · LEDGER

Scaling Vision Transformers

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2106.04560.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.04560 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:33:50.355434Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

16
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 14ddbe91-c3ed-47fe-9f7d-6bea1f796a44 · inbound

BEiT: BERT Pre-Training of Image Transformers cites this paper.

BEiT: BERT Pre-Training of Image Transformers Scaling Vision Transformers

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T11:50:11.512076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T11:50:11.476015Z digest=sha256:c3fe1adabb357f306ac689c66dc3d6a01bf567cc723f439fefb04d3702de2614

Observation 1153f87d-0441-4cdb-bf5a-47a5f824de54 · inbound

LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs cites this paper.

LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs Scaling Vision Transformers

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:01.108829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T10:21:01.062199Z digest=sha256:481d6b829b78582a9d714684076d6b9565ecf86ebbac4ec1fcfe3260ea0a133c

Observation e0b35262-610b-473b-b275-7e8eaa5a55e4 · inbound

Florence: A New Foundation Model for Computer Vision cites this paper.

Florence: A New Foundation Model for Computer Vision Scaling Vision Transformers

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:38:09.454541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:38:09.427509Z digest=sha256:f71dfa5bc442e29b50c240112b6bb7f99fd63c098212a12f9aec3cee3e900065

Observation 74d21f36-dd6b-4c4e-b66c-fdf1498937df · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning Scaling Vision Transformers

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.491031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:f520b8780b6e3bcab24b32d606c5d955e6286aacd4140af8c13651ce789a50ef

Observation b83230e8-7f76-4009-a14e-51caef88ada1 · inbound

LAION-5B: An open large-scale dataset for training next generation image-text models cites this paper.

LAION-5B: An open large-scale dataset for training next generation image-text models Scaling Vision Transformers

Reference 96

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:22:17.095075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T14:22:16.968028Z digest=sha256:832f2c1bd3f899d107e9836109c0ccf0fd23e69d87a65d5f697bb6bc13598068

Observation 614cd53c-7570-4844-80a3-6a848526af84 · inbound

LAION-5B: An open large-scale dataset for training next generation image-text models cites this paper.

LAION-5B: An open large-scale dataset for training next generation image-text models Scaling Vision Transformers

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:22:17.464938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T14:22:16.968028Z digest=sha256:2ecc0f22169b859ce6721a0a5e4ac16c03426441c72c784d6750f29cfb9109fa

Observation 92127002-de8e-4a05-a473-dd172a74731b · inbound

SemDeDup: Data-efficient learning at web-scale through semantic deduplication cites this paper.

SemDeDup: Data-efficient learning at web-scale through semantic deduplication Scaling Vision Transformers

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:43:30.924882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T02:43:30.851915Z digest=sha256:293a1fd36c42694c704a54edfa4c446f67377b3457d2657aa10e5a5c2b9c9e4b

Observation 4269597d-c48d-497b-b821-001dd2f2401f · inbound

Demystifying CLIP Data cites this paper.

Demystifying CLIP Data Scaling Vision Transformers

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:20:20.331713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T09:20:20.143143Z digest=sha256:0076636b1dda1748b22a77cc2cd8866c690a32d7de2aa6d611418cd3ba158b94

Observation fb1c1f3c-d5c6-4727-9985-6ac941003fa6 · inbound

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models cites this paper.

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models Scaling Vision Transformers

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-18T06:38:36.752601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T06:38:36.517935Z digest=sha256:1db3fe4a7ffa35614ce7303caf6e0032c99f5809b8ac3642d38fb62b24290168

Observation f4e70d8f-82f7-4799-bd4a-24d5a05dced0 · inbound

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model cites this paper.

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model Scaling Vision Transformers

Reference 251

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:30:03.073052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T17:30:02.803757Z digest=sha256:fc748bff335c8f2228058f04717f53952755fe4deb43ab003937431e15f0e174

Observation 6f83608e-8c5d-49e2-a675-91ffe946fce0 · inbound

Meek Models Shall Inherit the Earth cites this paper.

Meek Models Shall Inherit the Earth Scaling Vision Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:33:50.355434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:33:50.355434Z digest=sha256:d3db34e108f78874c8561630153e5a56d260f3496ec81e69b8925f2d86fc36a3

Observation 0163c443-6e58-4dce-bd00-d9e31354a8fb · inbound

Leveraging Transfer Learning and Mobile-enabled Convolutional Neural Networks for Improved Arabic Handwritten Character Recognition cites this paper.

Leveraging Transfer Learning and Mobile-enabled Convolutional Neural Networks for Improved Arabic Handwritten Character Recognition Scaling Vision Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T05:45:02.305660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:45:02.305660Z digest=sha256:08de3b7d2e971d5573fda983c08428adb836f98aaf63e809b5d64294c6efae10

Observation 53def06e-9750-43f1-9a49-ef25c4f0591a · inbound

SAM 3: Segment Anything with Concepts cites this paper.

SAM 3: Segment Anything with Concepts Scaling Vision Transformers

Reference 156

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.529795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:19e41c83b6d92789957e9fa39c794013248dfb6aa6ea8bef9fb132ee875dd6e2

Observation 4fb82683-6f09-4a56-9820-5139d5d1d7c3 · inbound

OmniMol: Transferring Particle Physics Knowledge to Molecular Dynamics with Point-Edge Transformers cites this paper.

OmniMol: Transferring Particle Physics Knowledge to Molecular Dynamics with Point-Edge Transformers Scaling Vision Transformers

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:20:57.963935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T13:18:02.842399Z digest=sha256:fa457fe427e3682233b179bc748dc27ede8fa8634812a2ab146a585c86ae67a0

Observation 8749f0a9-4f03-4ae0-8c9d-f691bb07e530 · inbound

Physics-Informed Transformer for Real-Time High-Fidelity Topology Optimization cites this paper.

Physics-Informed Transformer for Real-Time High-Fidelity Topology Optimization Scaling Vision Transformers

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:23:02.880853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T17:18:39.261579Z digest=sha256:ff955ae7c92ad740dcc4311b1b75aff0da17990242ae8de5eff466fab8250a66

Observation 9df39a8b-521d-4c8b-b3df-1ba46dfd9ab9 · inbound

Masked Contrastive Pre-Training Improves Music Audio Key Detection cites this paper.

Masked Contrastive Pre-Training Improves Music Audio Key Detection Scaling Vision Transformers

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:02.024344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:24:33.210955Z digest=sha256:bed9e16dd463171a501aa0c75b98a87b860fcd914bfd55348ea7e5b3f5f1e81c

Observation a558aaa5-e02f-41ac-a073-96bfc599dbd7 · inbound

Threats to Arabic Handwriting Recognition: Investigating Black-Box Adversarial Attacks on embedded ConvNet models cites this paper.

Threats to Arabic Handwriting Recognition: Investigating Black-Box Adversarial Attacks on embedded ConvNet models Scaling Vision Transformers

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:38:14.686617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:34:36.341683Z digest=sha256:a63d994ddfd319bf8691957d1db85c921212b99cf4f04c45fb9be73ae4b6054b

Observation cf687ffc-b345-4411-81ff-4f67edc3cdd4 · inbound

Empirical Bayes Conformal Prediction for Vision and Language Models cites this paper.

Empirical Bayes Conformal Prediction for Vision and Language Models Scaling Vision Transformers

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:45:19.791831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T04:45:11.816420Z digest=sha256:4dfe8b9c7505a4278a09e55d50ac3e9e6cba3d90a6e04126b8eba99340eeb0b9

Observation 96e3459a-788e-4879-b74a-317410fee9eb · inbound

Unified Neural Scaling Laws cites this paper.

Unified Neural Scaling Laws Scaling Vision Transformers

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:44:03.336396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T22:56:43.393302Z digest=sha256:1957ad775750d57bbcf5f425c4e627203cd3d219df3a4644d7d5a45891680fd6

Observation 6d85e87a-bf31-43d4-ab60-53873f257dde · inbound

Scaling Laws for Neural-Network Quantum States cites this paper.

Scaling Laws for Neural-Network Quantum States Scaling Vision Transformers

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:46:27.100672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T11:30:23.893472Z digest=sha256:b041524a44ae586ef4afbe0527ee00ac175f944cae4c188dbe948a5427a0a264

Observation e394515f-5dc5-40e8-acf1-0eaf6969930e · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Scaling Vision Transformers

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:30:07.801562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-25T20:05:09.179627Z digest=sha256:ee9df67286709a9403d51d0d10b0f7e6b8a54da946405a04c70bf4ae5fab30e7

Observation 52cbf837-3f41-4fe9-a848-0d1845e5fc5d · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Scaling Vision Transformers

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-02T10:14:13.365487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:14:13.365487Z digest=sha256:6447a32a352ed139dc195481adf7fbbe6197023de74ae86ab108f53e4da964fa

Observation 55b13f6c-2c76-4e87-b39a-92a8567d0e1e · inbound

Scale Weight Decay and Train Better cites this paper.

Scale Weight Decay and Train Better Scaling Vision Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.706869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.706869Z digest=sha256:f948fef2504bfc5af9d56fd3c4a4ee5372c525152008290b85089ec9aa5dcd79