Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:09:33.397262Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 0 inbound Pith citation observations for arXiv:2506.03433.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:09:33.397262Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
96 of 96 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3259072f-c0c8-465c-8dad-8d1ea6d68f7c · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Attention attention everywhere: Monocular depth prediction with skip attention
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a6dff74-e5a1-4030-bc6e-f33e08baecd7 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Self-supervised learning from images with a joint-embedding predictive architecture
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e73c98c1-debc-4e24-b8ee-4e211477e1ae · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Foundational Models Defining a New Era in Vision: A Survey and Outlook
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ba129c2-f5bd-4a00-9077-29881ed2d33c · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b5833a3-8193-4182-a5db-69f97e246698 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Beit: Bert pre-training of image transformers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dc0cbb8-fa7a-487f-899f-9edc42513260 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Adabins: Depth estimation using adaptive bins
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 015458de-a05f-44a2-b47e-cb13db6d2f52 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Language Models are Few-Shot Learners
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c79ea6be-af6d-4428-ae06-59686630daef · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Emerg- ing properties in self-supervised vision transformers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1cfa443-a58f-4865-9e6e-c8e2b946b5b0 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads MMDetection: Open MMLab Detection Toolbox and Benchmark
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ae5764e-d0bb-4e12-88dd-7b3ebd393ebb · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a51fa6-5d03-4d66-a720-74aeed9eb11b · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Mixformer: Mixing features across windows and dimensions
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1627e6d-8d87-4f4f-b0a3-da425aa2dc82 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Adaptformer: Adapt- ing vision transformers for scalable visual recognition
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f29d0633-c3aa-42a1-b7ef-8b85eda17217 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads A simple framework for contrastive learning of visual representations
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a9c32a1-630e-43b9-80ce-35eefb827812 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Vision transformer adapter for dense predictions
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ec0a4e6-e918-451e-907c-fc9cf16c620e · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b000884-592d-44fe-8e69-dfa7b8495789 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Masked-attention mask transformer for universal image segmentation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be2d17de-c24f-42a2-8a14-b95afb121481 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e461f47-cf2c-4408-8883-c89bdcf751d9 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Twins: Revisiting the design of spatial attention in vision transformers
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22d2b970-7161-4f76-a774-3ebcb701146b · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Mmsegmentation: Open- mmlab semantic segmentation toolbox and benchmark,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89ba021c-d6f1-44de-9e27-7679cb9af83c · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads The cityscapes dataset for semantic urban scene understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1dfef07-8d21-4eb9-89d1-5487746f385a · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Deformable convolutional networks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97a62938-dd01-4df9-9635-99b541c44670 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Instructblip: Towards general- purpose vision-language models with instruction tuning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b61b55ad-462f-47c2-84d7-a8c558f6347e · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Scaling vision transformers to 22 billion pa- rameters
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7d531c04-b57e-41e6-bd97-e66b886870ac · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads An image is worth 16x16 words: Trans- formers for image recognition at scale
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2fb3560-fa9d-430a-9bf8-7d30682d6726 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Eva: Exploring the limits of masked visual representa- tion learning at scale
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e99d77dd-cc26-4d5e-8413-36dbbd58fb33 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Deep ordinal regression net- work for monocular depth estimation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3fc27194-76cc-4bc4-ba45-1e646738c68a · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a9e5e659-6a19-4b56-947b-115dfb8a7e96 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Bootstrap your own latent-a new approach to self-supervised learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45015805-4052-4acf-9c11-d527202d93be · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads A survey on self-supervised learning: Algorithms, applications, and future trends
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f47fa069-4e62-48af-a500-1b9cf28921cd · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Vizwiz grand challenge: Answering visual questions from blind people
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7809092b-2956-4525-a06b-6afbddc61f86 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Flatten transformer: Vision transformer using fo- cused linear attention
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5edb397e-3521-460d-8699-efba01735197 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Deep residual learning for image recognition
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efb7723a-88fc-4cd1-965a-18ab8b0b525c · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Mask r-cnn
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2983936a-1c41-4978-9176-afedc02cfe0d · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Momentum contrast for unsupervised visual rep- resentation learning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1bbbe33b-d901-4c99-aee8-ec89037e9312 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Masked autoencoders are scalable vision learners
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fd28fc32-2224-4c36-85e3-d8335af63ef9 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Lora: Low-rank adaptation of large language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b2a137e0-97db-425e-84c9-ef78873510ca · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Introducing idefics: An open reproduction of state-of-the-art visual language model
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 442c063a-cfd3-441f-abc8-4d0d5d5b8b4b · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Oneformer: One transformer to rule universal image segmentation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4b8870aa-0c84-4aab-8234-897454c6b332 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Scaling up visual and vision-language representation learning with noisy text supervision
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 38f5c9eb-7431-4405-9d1b-6e7d0b2c5bdd · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Vi- sual prompt tuning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 967c7d6c-dad3-4300-b192-e839dd0a1f87 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Convolutional Bypasses Are Better Vision Transformer Adapters
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 530a8c4c-5fee-4b4a-b166-753686497d3c · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Fact: Factor-tuning for lightweight adaptation on vision transformer
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6824f71b-b31d-459c-912f-fdc7ad2c62c4 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Segment any- thing
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 38eac8ca-5c8a-400b-bc4a-c0c6258cc369 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Similarity of neural network representa- tions revisited
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7ea0415c-00f6-4ded-b754-5d323829846a · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Mask dino: Towards a unified transformer-based framework for object detection and segmentation
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 940846d3-e1ca-4e71-87c3-9ec17420200d · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 23514c34-1a8e-4d4a-9c7e-d94883a6abcd · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Benchmarking Detection Transfer Learning with Vision Transformers
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32c03cab-0be8-4dfe-a5bf-80fc4e7dc4f2 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Exploring plain vision transformer backbones for object de- tection
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3c753889-975f-4647-b694-b46ae19bb292 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Evaluating Object Hallucination in Large Vision-Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2bfbacc-3d5b-48d5-a79f-f8303ca9cd2c · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Visual Large Language Models for Generalized and Specialized Applications
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aedae50e-d8c5-429c-9a71-1036f9d57010 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Binsformer: Revisiting adaptive bins for monocular depth estimation
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 221cc480-74af-484c-a42a-739ca599f18c · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Microsoft coco: Common objects in context
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 402d1373-bcb3-4db1-be39-1f37bdff4500 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Va-depthnet: A variational approach to single image depth prediction
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b768bfb2-ceee-42df-8090-c5d879a4a74d · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Improved baselines with visual instruction tuning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6edd2ff8-1d5b-47dd-a504-0b6a6778b6a7 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Visual instruction tuning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1e7808ca-2b1e-4969-aa9b-c350a195dae5 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233, 2024
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1892988f-deeb-4a54-b3c8-3f6bf66b9eb5 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Swin transformer: Hierarchical vision transformer using shifted windows
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d909d31e-cc09-41ba-aa0e-caadb06e847b · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Swin transformer v2: Scaling up capacity and resolution
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6ed1515f-7dfb-427c-ab4c-d713d70df107 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads A convnet for the 2020s
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ce4e11f-e25b-4838-a428-f1eabc752a1f · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Decoupled Weight Decay Regularization
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6311134e-0259-4838-a3c6-dba2c3df9297 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1e8725e5-96fc-4b70-a576-2f353482bd17 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Time-memory-and parameter-efficient visual adaptation
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3da4c4c1-c22e-4f7e-bf3f-9712bfd7926e · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads The role of context for object detection and se- mantic segmentation in the wild
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 768dd031-b549-405a-821a-9093e571fb85 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads All in tokens: Uni- fying output space of visual tasks via soft token
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a8e34fd1-ed0c-4efd-9b2b-ac77a0aaffaf · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Dinov2: Learning robust visual features without super- vision
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1d470de2-7de9-4c78-931c-8d2c1eddd2f3 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads St-adapter: Parameter-efficient image-to-video transfer learning
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation eaaad0c8-d6bf-4eb8-8eb3-fd02189d1822 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads P3depth: Monocular depth estimation with a piecewise planarity prior
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7f985a76-b600-400d-92e2-203cf90d4391 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads idisc: Internal discretization for monocular depth estimation
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 079de25f-fd81-44e1-b743-8dbb0879aa80 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Learn- ing transferable visual models from natural language super- vision
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c5e79113-f7c9-476e-9772-452b273caa55 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Do vision trans- formers see like convolutional neural networks? NeurIPS, 34:12116–12128, 2021
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 34c1f187-4247-4add-8273-8f50d6476cfa · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Vi- sion transformers for dense prediction
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0f13137d-2af3-4c93-a9ce-fde997b11f10 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Learning multiple visual domains with residual adapters
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9e96e077-9e8c-4457-8363-09a34de5ffa2 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Iebins: Iterative elastic bins for monocular depth estimation
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a70ce1e0-d841-45de-94e2-4d38e66e0718 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Indoor segmentation and support inference from rgbd images
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8b5ff976-2a4b-4782-bbcd-1614c3ba9216 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Efficientnet: Rethinking model scaling for convolutional neural networks
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2ae3adb6-4b20-497d-9bf6-907e3b866b5e · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Training data-efficient image transformers & distillation through at- tention
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ebe46b12-e275-45aa-aa67-102b073baf57 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Pvt v2: Improved baselines with pyramid vision transformer
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0c88edec-f5a5-4c64-b025-35746b277c91 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Internimage: Exploring large-scale vi- sion foundation models with deformable convolutions
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 287062d4-a3bf-4e7c-a07b-b40b8cf97fff · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Vit-comer: Vision transformer with convolu- tional multi-scale feature interaction for dense predictions
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 63817e5e-53f0-4dbe-8191-3fba7dda9178 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Unified perceptual parsing for scene understand- ing
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8c4a6ed8-c6a8-4050-b989-c8a690177d4d · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Focal Self-attention for Local-Global Interactions in Vision Transformers
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0ab7dbb-5722-43e5-b715-e735e7f345e9 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Depth anything: Unleashing the power of large-scale unlabeled data
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3c4781a8-e821-41d6-b17d-c6366900841e · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Visual tuning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ee55490a-26a2-4b93-b9fa-1deade12af1d · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Coca: Contrastive captioners are image-text foundation models
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d3f3e392-39d7-424b-ac60-5e27922ff185 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 152b1732-5b13-47fc-9073-523dd585f361 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Neural window fully-connected crfs for monocular depth estimation
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e91391c4-2d1e-4b2f-8e56-80ab8e863f47 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Spanet: Frequency-balancing token mixer using spectral pooling aggregation modulation
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 722a9709-a5b9-4c82-934d-07082b6f963f · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a97d7428-4760-4ec0-ac4e-0d40505cc666 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Sigmoid loss for language image pre-training
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5f21ecd9-32a9-4f1f-8eda-75b0b2becc75 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Memory efficient transformer adapter for dense pre- dictions
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9add61c0-1e67-4eb3-9aac-443fa919854b · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads A Survey of Large Language Models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2d4daee-3c0d-4ae2-9304-95c7e4902b2a · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Semantic under- standing of scenes through the ade20k dataset
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 63af6f46-6ab1-4d00-8c30-9cb4d3a1d2cf · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads A Comprehensive Survey on Pretrained Foundation Models: A History from BERT to ChatGPT
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6730720-017b-4f8a-8bd4-88f3cc8a04c2 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Image bert pre-training with online tokenizer
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8f1f3dac-08b2-47be-ac72-112bfc71bfd1 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Conditional prompt learning for vision-language models
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d994396c-d3b9-46b3-9c55-d3ef05e317a6 · outbound
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Learning to prompt for vision-language models
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
No inbound Pith citation observations are available.