REVIEW 2 major objections 5 minor 57 references
Heterogeneous vision foundation models become reliably stitchable with final-feature matching and can beat either parent model.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 22:18 UTC pith:IFQ6C3GD
load-bearing objection Solid empirical recipe: FFM makes heterogeneous VFMs stitchable, and VST turns that into a real multi-backbone efficiency knob. the 2 major comments →
Revisiting Model Stitching In the Foundation Model Era
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
With a final-feature-matching loss at the target’s penultimate layer followed by task fine-tuning, heterogeneous vision foundation models become reliably stitchable across vision tasks; at deep stitch points the stitched model can surpass either constituent model while adding only the small overhead of the stitch layer, and the gains over matched self-stitch controls indicate genuine complementary fusion rather than extra capacity alone.
What carries the argument
Final Feature Matching (FFM): train only the stitch layer so the frozen target’s final output features are reproduced by the stitched path, then fine-tune with the task loss. Paired with self-stitch controls that insert the identical module into a single model, this two-stage procedure is what makes stitching reliable and grounds the VFM Stitch Tree that shares early layers while keeping specialized deep branches.
Load-bearing premise
The self-stitch baseline—inserting the same trainable module into source-only or target-only models at the same depth, with the same losses and data—fully isolates capacity and task-adaptation effects, so any leftover gain must be genuine cross-model complementarity.
What would settle it
On held-out classification and segmentation tasks, train cross-VFM stitches and identical self-stitch baselines for several model pairs at multiple depths; if the cross-VFM stitch never significantly exceeds the better self-stitch, the claim that stitching fuses complementary strengths fails.
If this is right
- Multimodal systems that currently run several full vision foundation models can share early layers and keep only specialized deep branches, recovering a large fraction of the multi-backbone gain at far lower extra cost.
- Stitch layers can be pre-trained on task-agnostic image data and reused, enabling hybrid foundation models without per-task retraining of the connector.
- Deep stitches can improve accuracy over either parent model while adding only a light projector at inference.
- Success or failure of stitching at different depths marks where the representations of differently trained models align or diverge.
Where Pith is reading between the lines
- The same final-feature-matching recipe is a candidate for stitching language or audio foundation models that also differ in data and objectives.
- If early layers are largely pretraining-specific, systems may gain more by discarding or heavily adapting them than by forcing a stitch at shallow depth.
- Stitch Trees suggest a general design pattern for any stack that currently loads multiple frozen backbones side-by-side.
- Systematic failure when a weak model is the source supplies a practical test of whether an encoder’s intermediate features still retain recoverable task-critical information.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper revisits model stitching for heterogeneous Vision Foundation Models (VFMs) that differ in objectives, data, and modality (CLIP, DINOv2, DINOv3, SigLIP2). It shows that conventional Layer Feature Matching (LFM) and naive Task-Loss Training (TLT) fail, especially at shallow stitch points, while a two-stage recipe—Final Feature Matching (FFM) at the target penultimate layer followed by task fine-tuning—makes VFMs reliably stitchable. Stitched models often match or exceed linear probes of both constituents and consistently outperform self-stitch controls across classification (fMoW, iNaturalist, Aircraft) and ADE20K segmentation. Building on this, the authors introduce VFM Stitch Tree (VST), which shares early layers and retains specialized deep layers, recovering a large fraction of multi-VFM gains in a MoF-LLaVA setting at far lower compute/memory cost.
Significance. If the results hold, the work converts stitching from a diagnostic probe into a practical integration tool for complementary VFMs. The controlled ablations (LFM vs FFM vs TLT, self-stitch baselines, multiple depths and stitch-layer families, four VFMs, classification plus dense prediction) are thorough and reproducible in spirit. The VST construction supplies a concrete accuracy–latency knob for multi-VFM multimodal LLMs, addressing a real deployment cost that scales linearly with the number of backbones. The absolute stitchability result (Findings 2–3) and the efficiency numbers remain useful even under residual caveats about self-stitch optimization.
major comments (2)
- Section 4.1–4.2 and Figure 5: the self-stitch baseline is the main control isolating capacity/task-adaptation from complementarity. While the paper shows consistent gains over self-stitch across depths, model pairs, and tasks (Tables 2, 7–11), it does not fully rule out architectural under-optimization of self-stitch (e.g., the stitch module may interact differently when source and target are identical). A short additional control—e.g., random-feature or identity-initialized self-stitch, or a capacity-matched residual adapter—would strengthen the fusion claim without changing the central stitchability result.
- Section 6 / Figure 9 / Table 14: VST is evaluated only on VQAv2 and MME (Perception/Cognition) inside a single MoF-LLaVA (CLIP+DINOv2) setup with Qwen-3B. The efficiency claims (4.3 % / 39 % extra resources recovering 45 % / 84 % of the two-VFM gain) are promising, but the multimodal evidence is narrower than the vision-only evidence. Expanding to at least one additional MLLM benchmark or a second multi-VFM pair would make the application claim more robust.
minor comments (5)
- Figure 2 caption and surrounding text: the symbols for layer vs final feature distance are described but not rendered consistently in the text; a short legend or explicit marker names would help.
- Table 3: LoRA underperforms MLP despite higher expressiveness; a one-sentence hypothesis is given, but a brief ablation (rank, which layers receive LoRA) would clarify whether the result is capacity- or optimization-related.
- Section 5.4 / Table 4: task-agnostic FFM on LLaVA-1.5 data is interesting; stating the exact stitch position and whether any task-specific TLT is applied after would avoid ambiguity.
- Appendix A.1: input resolutions and patch sizes differ across VFMs (336/14 vs 384/16); a short note on how token counts are aligned at the stitch layer would improve reproducibility.
- Minor typos: “na¨ıve” rendering, occasional missing spaces around arrows (DINOv2→SigLIP2), and “SigLIP 2” vs “SigLIP2” inconsistency.
Circularity Check
No significant circularity: purely empirical protocol with held-out benchmarks and independent self-stitch controls; no derivation reduces claimed gains to fitted inputs or self-citation by construction.
full rationale
The paper's central claims (heterogeneous VFMs become stitchable via Final Feature Matching + task fine-tuning; stitched models can exceed self-stitch and linear-probe baselines; VST yields controllable multi-VFM efficiency) rest entirely on experimental measurements. Stitch layers are trained on labeled or unlabeled training splits (fMoW, iNaturalist, Aircraft, ADE20K, LLaVA-1.5 data) and evaluated on held-out test sets or standard MLLM benchmarks (VQAv2, MME). The self-stitch baseline inserts an identical trainable module into source-only or target-only models under the same losses, positions, and data, providing an independent capacity/task-adaptation control rather than a definitional identity. No equation equates a reported accuracy gain to a fitted parameter by construction; no uniqueness theorem or load-bearing premise is imported solely via overlapping-author citation; prior stitching literature is cited for historical context, not as an unverified axiom that forces the present results. The work is therefore self-contained against external benchmarks and free of the enumerated circularity patterns.
Axiom & Free-Parameter Ledger
free parameters (3)
- stitch-layer learning rates =
{0.001, 0.005, 0.01}
- stitch positions =
2/6/10/14/18/22
- MLP stitch architecture (hidden size, ReLU) =
two-layer ReLU MLP
axioms (3)
- domain assumption Source and target VFM weights remain completely frozen; only the stitch layer is trainable.
- domain assumption Linear probing (or a linear decoder for segmentation) on frozen final features is a valid proxy for representational quality and task performance.
- domain assumption L2 distance between feature maps is a suitable matching objective for both layer-wise and final-feature alignment.
invented entities (2)
-
VFM Stitch Tree (VST)
no independent evidence
-
Final Feature Matching (FFM) training recipe
no independent evidence
read the original abstract
Model stitching, connecting early layers of one model (source) to later layers of another (target) via a light stitch layer, has served as a probe of representational compatibility. Prior work finds that models trained on the same dataset remain stitchable (negligible accuracy drop) despite different initializations or objectives. We revisit stitching for Vision Foundation Models (VFMs) that vary in objectives, data, and modality mix (e.g., CLIP, DINOv2, SigLIP 2) and ask: Are heterogeneous VFMs stitchable? We introduce a systematic protocol spanning the stitch points, stitch layer families, training losses, and downstream tasks. Three findings emerge. (1) Stitch layer training matters: conventional approaches that match the intermediate features at the stitch point or optimize the task loss end-to-end struggle to retain accuracy, especially at shallow stitch points. (2) With a simple feature-matching loss at the target model's penultimate layer, heterogeneous VFMs become reliably stitchable across vision tasks. (3) For deep stitch points, the stitched model can surpass either constituent model at only a small inference overhead (for the stitch layer). Building on these findings, we further propose the VFM Stitch Tree (VST), which shares early layers across VFMs while retaining their later layers, yielding a controllable accuracy-latency trade-off for multimodal LLMs that often leverage multiple VFMs. Taken together, our study elevates stitching from a diagnostic probe to a practical recipe for integrating complementary VFM strengths and pinpointing where their representations align or diverge.
Reference graph
Works this paper leans on
-
[1]
Functional similarity by functional latent alignment
Ioannis Athanasiadis, Anmar Karmush, and Michael Fels- berg. Functional similarity by functional latent alignment. InGreeks in AI Symposium 2025, 2025. 3
2025
-
[2]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. InarXiv preprint arXiv:2309.16609, 2023. 2
Pith/arXiv arXiv 2023
-
[3]
Revisit- ing model stitching to compare neural representations
Yamini Bansal, Preetum Nakkiran, and Boaz Barak. Revisit- ing model stitching to compare neural representations. InAd- vances in Neural Information Processing Systems (NeurIPS),
-
[4]
Emerg- ing properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 9650–9660, 2021. 1
2021
-
[5]
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Ge- offrey Hinton. A simple framework for contrastive learning of visual representations. InInternational conference on ma- chine learning, pages 1597–1607. PmLR, 2020. 4
2020
-
[6]
Pali: A jointly-scaled multilingual language-image model.arXiv preprint arXiv:2209.06794, 2022
Xi Chen, Xiao Wang, Soravit Changpinyo, Anthony J Pier- giovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. Pali: A jointly-scaled multilingual language-image model.arXiv preprint arXiv:2209.06794, 2022. 3
Pith/arXiv arXiv 2022
-
[7]
Functional map of the world.Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), 2018
Gordon Christie, Neil Fendley, James Wilson, and Ryan Mukherjee. Functional map of the world.Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), 2018. 4, 6, 1
2018
-
[8]
How not to stitch representations to measure similarity: Task loss matching versus direct matching
Katherine M Collins, Umang Bhatt, and Adrian Weller. How not to stitch representations to measure similarity: Task loss matching versus direct matching. InProceedings of the AAAI Conference on Artificial Intelligence, 2025. 2, 5
2025
-
[9]
Matszan- gosz, Gergely Papp, and D´aniel Varga
Adri ´an Csisz ´arik, P ´eter K ˝or¨osi-Szab´o, ´Akos K. Matszan- gosz, Gergely Papp, and D´aniel Varga. Similarity and match- ing of neural network representations. InAdvances in Neu- ral Information Processing Systems (NeurIPS), pages 5656– 5668, 2021. 2, 3
2021
-
[10]
Data fil- tering networks.arXiv preprint arXiv:2309.17425, 2023
Alex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt, Alexander Toshev, and Vaishaal Shankar. Data fil- tering networks.arXiv preprint arXiv:2309.17425, 2023. 3
Pith/arXiv arXiv 2023
-
[11]
Multi- modal autoregressive pre-training of large vision encoders
Enrico Fini, Mustafa Shukor, Xiujun Li, Philipp Dufter, Michal Klein, David Haldimann, Sai Aitharaju, Victor G Turrisi da Costa, Louis B ´ethune, Zhe Gan, et al. Multi- modal autoregressive pre-training of large vision encoders. InProceedings of the Computer Vision and Pattern Recogni- tion Conference, pages 9641–9654, 2025. 4
2025
-
[12]
Mme: A comprehensive evaluation benchmark for multimodal large language models
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, et al. Mme: A comprehensive evaluation benchmark for multimodal large language models. InarXiv preprint arXiv:2306.13394, 2023. 1, 2
Pith/arXiv arXiv 2023
-
[13]
Mme: A comprehensive evaluation bench- mark for multimodal large language models
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, et al. Mme: A comprehensive evaluation bench- mark for multimodal large language models. InThe Thirty- ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2025. 8
2025
-
[14]
Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Ba- tra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 6904–6913, 2017. 8, 2
2017
-
[15]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. InProceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016. 2
2016
-
[16]
Lm-deeplabv3+: A lightweight image segmentation algorithm based on multi- scale feature interaction.Applied Sciences, 14(4):1558,
Xinyu Hou, Peng Chen, and Haishuo Gu. Lm-deeplabv3+: A lightweight image segmentation algorithm based on multi- scale feature interaction.Applied Sciences, 14(4):1558,
-
[17]
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arxiv 2021.arXiv preprint arXiv:2106.09685, 10, 2021. 6
Pith/arXiv arXiv 2021
-
[18]
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Liang Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models. Iclr, 1(2):3, 2022. 6
2022
-
[19]
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. InProceedings of the IEEE/CVF con- ference on computer vision and pattern recognition, pages 6700–6709, 2019. 2
2019
-
[20]
Binyuan Hui, Jian Yang, Zeyu Cui, Jiaxi Yang, Dayiheng Liu, Lei Zhang, Tianyu Liu, Jiajun Zhang, Bowen Yu, Kem- ing Lu, et al. Qwen2. 5-coder technical report.arXiv preprint arXiv:2409.12186, 2024. 8
Pith/arXiv arXiv 2024
-
[21]
Openclip.https://github.com/mlfoundations/ open_clip, 2021
Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, et al. Openclip.https://github.com/mlfoundations/ open_clip, 2021. 2, 3
2021
-
[22]
Benchmarking the robustness of semantic segmentation models
Christoph Kamann and Carsten Rother. Benchmarking the robustness of semantic segmentation models. InProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 6
2020
-
[23]
How to benchmark vision foundation models for semantic seg- mentation? InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1162– 1171, 2024
Tommie Kerssies, Daan De Geus, and Gijs Dubbelman. How to benchmark vision foundation models for semantic seg- mentation? InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1162– 1171, 2024. 6, 2
2024
-
[24]
Openvla: An open-source vision-language-action model.arXiv preprint arXiv:2406.09246, 2024
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al. Openvla: An open-source vision-language-action model.arXiv preprint arXiv:2406.09246, 2024. 7
Pith/arXiv arXiv 2024
-
[25]
Similarity of neural network models revisited: Measuring functional similarity
Max Klabunde, Tobias Schubert, and Sebastian Lapuschkin. Similarity of neural network models revisited: Measuring functional similarity. InInternational Conference on Ma- chine Learning (ICML), 2024. 2
2024
-
[26]
Similarity of neural network represen- tations revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton. Similarity of neural network represen- tations revisited. InInternational conference on machine learning, pages 3519–3529. PMlR, 2019. 2
2019
-
[27]
Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123(1):32–73, 2017
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalan- tidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123(1):32–73, 2017. 2
2017
-
[28]
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 2
2009
-
[29]
Understanding image repre- sentations by measuring their equivariance and equivalence
Karel Lenc and Andrea Vedaldi. Understanding image repre- sentations by measuring their equivariance and equivalence. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015. 2
2015
-
[30]
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Doll ´ar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 936–944, 2017. 6
2017
-
[31]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. InAdvances in Neural Information Processing Systems (NeurIPS), 2023. 2
2023
-
[32]
Improved baselines with visual instruction tuning.arXiv preprint arXiv:2310.03744, 2024
Haotian Liu, Chunyuan Li, Yong Jae Li, and Yong Jae Lee. Improved baselines with visual instruction tuning.arXiv preprint arXiv:2310.03744, 2024. 4, 7, 8, 2
Pith/arXiv arXiv 2024
-
[33]
Fine-tuning is fine, if cal- ibrated.Advances in neural information processing systems, 37:136084–136119, 2024
Zheda Mai, Arpita Chowdhury, Ping Zhang, Cheng-Hao Tu, Hong-You Chen, Vardaan Pahuja, Tanya Berger-Wolf, Song Gao, Charles Stewart, Yu Su, et al. Fine-tuning is fine, if cal- ibrated.Advances in neural information processing systems, 37:136084–136119, 2024. 1
2024
-
[34]
Zheda Mai, Arpita Chowdhury, Zihe Wang, Sooyoung Jeon, Lemeng Wang, Jiacheng Hou, and Wei-Lun Chao. Ava- bench: Atomic visual ability benchmark for vision founda- tion models.arXiv preprint arXiv:2506.09082, 2025. 3
Pith/arXiv arXiv 2025
-
[35]
Lessons and insights from a unifying study of parameter-efficient fine-tuning (peft) in visual recognition
Zheda Mai, Ping Zhang, Cheng-Hao Tu, Hong-You Chen, Quang-Huy Nguyen, Li Zhang, and Wei-Lun Chao. Lessons and insights from a unifying study of parameter-efficient fine-tuning (peft) in visual recognition. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 14845–14857, 2025. 1, 6
2025
-
[36]
Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew Blaschko, and Andrea Vedaldi. Fine-grained visual classi- fication of aircraft.http://www.robots.ox.ac.uk/ ˜vgg/data/fgvc- aircraft/, 2013. arXiv preprint arXiv:1306.5151. 6, 1
Pith/arXiv arXiv 2013
-
[37]
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timoth ´ee Darcet, Th´eo Moutakanni, Huy V V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 2, 3, 1
Pith/arXiv arXiv 2023
-
[38]
Stitchable neural networks
Zizheng Pan, Bohan Zhuang, Haoyu He, Jing Liu, and Jian- fei Cai. Stitchable neural networks. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16041–16050, 2023. 3
2023
-
[39]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. InProceedings of the 38th International Conference on Machine Learning (ICML), 2021. 1, 2, 3
2021
-
[40]
Oriane Sim ´eoni, Huy V . V o, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Micha ¨el Ramamonjisoa, Francisco Massa, Daniel Haziza, Luca Wehrstedt, Jianyuan Wang, Timoth´ee Darcet, Th´eo Moutakanni, Leonel Sentana, Claire Roberts, John Brandt, Camille Couprie, Julien Mairal, Herv´e J ´ego...
Pith/arXiv arXiv 2025
-
[41]
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards vqa models that can read. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8317–8326, 2019. 2
2019
-
[42]
Functional alignment can mislead: Examining model stitch- ing
Damian Smith, Harvey Mannering, and Antonia Marcu. Functional alignment can mislead: Examining model stitch- ing. InProceedings of the 42nd International Conference on Machine Learning (ICML), 2025. 2
2025
-
[43]
Eva-clip: Improved training techniques for clip at scale
Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. Eva-clip: Improved training techniques for clip at scale. arXiv preprint arXiv:2303.15389, 2023. 3
Pith/arXiv arXiv 2023
-
[44]
Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024
Peter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Adithya Jairam Vedagiri IYER, Sai Charitha Akula, Shusheng Yang, Jihan Yang, Manoj Middepogu, Ziteng Wang, et al. Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024. 7, 4
2024
-
[45]
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie. Eyes wide shut? exploring the visual shortcomings of multimodal llms.arXiv preprint arXiv:2401.06209, 2024. 1, 3, 7, 2
Pith/arXiv arXiv 2024
-
[46]
Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muham- mad Ferjad Naeem, Ibrahim Alabdulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, et al. Siglip 2: Multilingual vision-language en- coders with improved semantic understanding, localization, and dense features.arXiv preprint arXiv:2502.14786, 2025. 3, 1, 4
Pith/arXiv arXiv 2025
-
[47]
The inaturalist species classification and detection dataset
Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge Belongie. The inaturalist species classification and detection dataset. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 8769–8778, 2018. 6, 1
2018
-
[48]
Deep model reassembly.Advances in neural information processing systems, 35:25739–25753, 2022
Xingyi Yang, Daquan Zhou, Songhua Liu, Jingwen Ye, and Xinchao Wang. Deep model reassembly.Advances in neural information processing systems, 35:25739–25753, 2022. 3
2022
-
[49]
Sigmoid loss for language-image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language-image pre-training. arXiv preprint arXiv:2303.15343, 2023. 1, 2, 3
Pith/arXiv arXiv 2023
-
[50]
Scene parsing through ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. InProceedings of the IEEE Conference on Computer Vision and Pattern Vision (CVPR), 2017. 6, 1
2017
-
[51]
Semantic under- standing of scenes through the ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fi- dler, Adela Barriuso, and Antonio Torralba. Semantic under- standing of scenes through the ade20k dataset. InInterna- tional Journal of Computer Vision, 2019. 6, 1
2019
-
[52]
Baichuan Zhou, Ying Hu, Xi Weng, Junlong Jia, Jie Luo, Xien Liu, Ji Wu, and Lei Huang. Tinyllava: A frame- work of small-scale large multimodal models.arXiv preprint arXiv:2402.14289, 2024. 2 Revisiting Model Stitching in the Foundation Model Era Supplementary Material We provide details omitted in the paper. • Sec. A: Experiment and dataset details • Sec...
Pith/arXiv arXiv 2024
-
[53]
Training is performed with automatic mixed precision usingbfloat16. A.2.3. Semantic Segmentation A.2.4. Dataset Details ADE20K [50, 51]ADE20K is a scene-centric seman- tic segmentation dataset with pixel-level annotations for 150 object and stuff categories across diverse indoor and outdoor environments. We adopt the canonical split with 20,210 training i...
-
[54]
Full” denotes running all VFMs independently. “VST-n
[32], we base our calculations on a 23-layer depth. While the “Full” setting incurs a 300% computational over- head compared to a single VFM, VST-14 significantly re- duces this burden. By sharing the first 14 layers and main- taining specialized branches only from layer 15 onwards, VST-14 requires processing only3×(23−14) = 27ad- ditional layers. This re...
-
[55]
Define the Upper Bound (Denominator) (A) CLIP Baseline 91.75 58.74 69.00 1418.5 277.1 - 0% (B) Full (CLIP + DINOv2) 92.72 61.64 70.30 1460.3 311.8 - 100% (C) Max Gain (∆max =B−A) 0.97 2.90 1.30 41.8 34.7 (Denom.) -
-
[56]
VST-22: The Lightweight Knob (High Efficiency) (D) VST-22 (Ours) 92.12 59.21 69.15 1451.6 305.7 -4.3% (E) VST Gain (∆VST =D−A) 0.37 0.47 0.15 33.1 28.6 (Num.) - Normalized Gain (% =E/C) 38.1% 16.2% 11.5% 79.1% 82.5% 45.5% 4.3%
-
[57]
Lightweight
VST-14: The Balanced Knob (High Performance) (F) VST-14 (Ours) 92.54 60.69 69.88 1474.4 301.8 - 39.0% (G) VST Gain (∆VST =F−A) 0.79 1.95 0.88 55.9 24.7 (Num.) - Normalized Gain (% =G/C) 81.4% 67.2% 67.7% 133.7% 71.1% 84.2%39.0% Table 14.Detailed Calculation of Normalized Gain.Comparison of the “Lightweight” VST-22 and the “Balanced” VST-14. The Orange Row...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.