REVIEW 3 major objections 5 minor 87 references
MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Interleaving visual object tokens into caption text during alignment pretraining gives a 12-fold data-efficiency gain over image-caption pretraining, matching or beating 600K pairs with 50K samples.
desk verdict A genuinely novel alignment mechanism with a well-controlled study, but the headline generalization claims are not yet supported because the evaluation images overlap the training corpus. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the code-switched sequence $X_{\mathrm{MMCS}}=\operatorname{Concat}(X_{<i}, v_{\mathrm{object}}, X_{>i+m})$: a caption in which each textual entity $e=[x_i,\dots,x_{i+m}]$ is replaced by the image tokens $v_{\mathrm{object}}$ whose spatial regions intersect the object's bounding box. The projector is trained with a language-modeling loss over the remaining text tokens plus an entity-reconstruction loss $\mathcal{L}_{\mathrm{entity}}=-\log p_\theta(e\mid X_{<i}^{\mathrm{MMCS}}, v_{\mathrm{object}})$, which forces visual object tokens to carry identity-bearing information capable of regenerating the original phrase. The data-synthesis pipeline (captioner, entity extractor, Grounding DINO localizer, SAM-2.1 segmenter, box/text thresholds 0.4/0.3, and the 20% mask-area rule) supplies the one-to-one correspondences that make the substitution meaningful.
What would settle it
Take the 50K MMCS pretraining set, randomly replace 10% of the object bounding boxes with unrelated same-size boxes from the same image, and train the identical Qwen2.5-3B setup; if that noised run no longer beats the 600K image-caption baseline on RefCOCO, RefCOCO+, and RefCOCOg, the data-efficiency claim depends on near-perfect synthetic correspondences rather than on the code-switching format itself.
Extended reading notes
Core claim
The paper claims that the standard practice of aligning a global image representation with a long caption leaves referential ambiguity—the model must guess which visual region corresponds to which textual phrase—and that this ambiguity is the root cause of data inefficiency and weak grounding in multimodal LLMs. MMCS removes the guesswork by mechanically interleaving: for each noun-phrase entity in the caption, the textual tokens are replaced by the image tokens of the corresponding object (Eq. 1), so next-token prediction and an explicit entity-reconstruction loss (Eq. 3) are conditioned directly on local visual features. Trained on 773K such synthetic code-switched samples, a Qwen2.5-3B-based model with only 50K samples matches or exceeds an image-caption baseline trained on 600K samples, and the full setup improves referring-expression grounding by an average of 7.9% and perception-centric benchmarks by 2.1% across three LLM backbones and two vision encoders. The paper supports the mechanism with layer-wise CKA/CKNNA/Mutual k-NN alignment scores and attention maps.
Load-bearing premise
The load-bearing premise is that the automated pipeline produces one-to-one object-entity correspondences accurate enough at scale to teach rather than mislead, and that the object-level grounding learned in interleaved pretraining still transfers when SFT and inference revert to global image representations.
Editorial extensions
If this is right
- A 50K-sample MMCS pretraining run outperforms a 600K-sample image-caption run on the same downstream SFT, so the data-efficiency gain is a factor of roughly 12, not an incremental improvement.
- Grounding gains hold across model scales: 3B and 8B LLMs with different vision encoders all improve on RefCOCO, RefCOCO+, and RefCOCOg, averaging +7.9%.
- The two loss terms have separable roles: dropping $\mathcal{L}_{\mathrm{entity}}$ costs 4.0 points on perception, while dropping the language-modeling loss costs 4.9 points on grounding; both are needed.
- The alignment learned during interleaved pretraining transfers to standard global-image SFT and inference, requiring no coordinate tokens or grounding modules at inference time.
- Representation-level evidence (CKA, CKNNA, Mutual k-NN) and sharper attention maps are consistent with the claim that explicit object-entity supervision, not merely more data, improves cross-modal alignment.
Reading between the lines
- A direct extension not explored in the paper: apply MMCS to domains where text regions are the objects—charts, documents, scene text—since the pipeline's measured implicit OCR supervision and the paper's own limitation statement suggest the same substitution mechanism would work there.
- The reported reduction of roughly 41% in image tokens during pretraining implies MMCS could lower alignment-stage compute cost as well as data cost; the paper reports the token count but does not model total FLOPs or energy.
- If the observed trend continues past 1M samples, object-level code-switching could combine naturally with bootstrapping loops in which the model proposes its own entity-grounding pairs, a path the paper lists as future work but does not test.
- The part-of-speech negative-log-likelihood results suggest the benefit is not confined to nouns: verbs and prepositions also become more predictable, hinting that object-level grounding could scaffold learning of relations and actions without explicit relation labels.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MultiModal Code-Switching (MMCS), a pretraining paradigm for multimodal LLMs that replaces textual entity tokens in captions with the corresponding visual object tokens, training with a language modeling loss and an entity reconstruction loss. A pipeline using Qwen-based captioning/entity extraction, Grounding DINO, and SAM-2.1 generates a 773K-sample pretraining dataset. Experiments report data efficiency (50K samples matching 600K image-level pretraining) and gains on visual grounding (+7.9%) and perception benchmarks, with mechanistic analyses via CKA/CKNNA/Mutual k-NN and attention maps.
Significance. If the gains persist on disjoint images, this is a meaningful improvement over image-level alignment, providing explicit object-level supervision during pretraining. The paper's controlled setup (identical image-caption pairs for MMCS and the image-level baseline) and robustness ablations (caption-source swap, correspondence perturbation, resolution sweep) are strengths. The release of code and dataset supports reproducibility.
major comments (3)
- [Section 4.1 / Appendix A.1 / A.3] The evaluation benchmarks overlap with the pretraining and SFT corpora. RefCOCO is built on COCO images, and COCO and GQA images are included in the pretraining sources (Appendix A.1); GQA is also in the SFT set (Section A.2, which explicitly lists GQA). The paper reports no overlap statistics and no evaluation on a disjoint image set. Consequently, the headline claims of data efficiency (Figure 4/Table 11) and transferable grounding gains (Table 1) could be inflated by memorized object layouts rather than a generalizable object-level alignment mechanism. Please either evaluate on benchmarks whose images are provably disjoint from all training data, or report results on the non-overlapping subsets of the current benchmarks, and discuss any performance difference.
- [Appendix A.1, Table 5; Section 3.3] The quality of the synthesized correspondences is validated on only 200 objects per source by a VLM judge and 20 objects per source by human annotators. This is a thin basis for the claim of 'accurate object-entity correspondences' at 773K scale, particularly because the filtering thresholds (box threshold 0.4, text threshold 0.3, mask-area fraction 0.2) are hand-chosen. Although the perturbation analysis (D.2) shows robustness to random box noise, it does not address systematic errors in entity extraction or captioning. Please provide a larger human-validated sample or a per-category error analysis, and report the sensitivity of downstream performance to the filtering thresholds.
- [Section 4.1; Section 5] The pretraining objective operates on interleaved sequences where textual entities are replaced by object tokens, yet SFT and inference revert to global image representations. The mechanistic evidence of improved alignment (Figure 5) and sharper attention (Figure 6) is presented for pretraining-only models, so it is not directly shown that these improvements survive the format mismatch after SFT, when the downstream numbers are measured. Please either report the alignment metrics and attention analyses after SFT or demonstrate via ablations (e.g., SFT with interleaved inputs for both methods) that the transfer path is consistent.
minor comments (5)
- [Section 4.1 and A.2] The dataset name is mistyped as 'LLaV A-NeXT' twice; it should be 'LLaVA-NeXT'.
- [Table 2 and Table 14] The abbreviation 'VQAT' is used without definition; it should be 'TextVQA'.
- [Section 5.2] Specify whether the attention maps are from pretraining-only checkpoints or after SFT, to avoid ambiguity.
- [Figure 7 caption] State the method for computing the 41.4% reduction explicitly (e.g., average token count across 3,000 samples per source).
- [Section D.5] The paired t-test is only reported for general VQA; consider reporting per-benchmark significance or confidence intervals for the grounding and perception gains.
Circularity Check
No circularity: the MMCS objective and its controlled baselines are self-contained; reported evaluation overlaps are generalization risks, not circular reductions.
full rationale
The paper's central derivation chain is not circular. The MMCS pretraining objective (Eq. 4, combining L_LM and L_entity) is trained on a synthesized dataset of 773K samples and evaluated on external benchmarks such as RefCOCO, OCRBench, CVBench, MMB, and GQA; no method constant or loss weight is fitted to evaluation outcomes. The central controlled comparison in Section 4.1 explicitly states that 'both MMCS and the image-level baseline are trained on identical image–caption pairs,' so the reported gains are attributable to the code-switching formulation rather than to data quality or to a fitted input. The robustness experiments in Appendix D.1 and D.2 further show the gains survive changes in caption source and injected 10–30% correspondence noise, indicating the result is not built into the objective by construction. The representation-alignment analyses (CKA, CKNNA, Mutual k-NN in Section 5.1 and Figure 5) are post-hoc measurements on separately pretrained checkpoints, not quantities defined in terms of downstream performance. The paper's self-citations are not load-bearing; references such as Xing et al. (2025) support the background claim that dense captions are common for alignment, and no uniqueness theorem or ansatz is imported from the authors' prior work. The Limitations section and the acknowledged uneven object-category coverage are honest scope caveats about generalization and data bias, not admissions of circularity. The skeptic's concern that pretraining/SFT corpora overlap evaluation images (COCO/GQA/TextVQA) is a benchmark-contamination and generalization risk, which is a correctness or evaluation-validity issue, not a circularity in the derivation of MMCS from its inputs. Under the hard rule that circularity must be exhibited by an explicit reduction of a claimed prediction to a fitted input or self-citation chain, no such step exists in this paper.
Assumptions & free parameters
free parameters (4)
- Grounding DINO confidence thresholds =
box 0.4, text 0.3
- SAM-2.1 mask-area filter =
20% of bounding-box area
- Bounding-box size constraints =
reject boxes smaller than one visual patch or larger than 50% of image area
- Relative weight of L_entity vs LLM =
1.0 (unweighted sum)
assumptions (4)
- domain assumption Image-level alignment suffers from referential ambiguity that limits grounding and data efficiency
- domain assumption Image tokens whose spatial regions intersect an object's bounding box form a sufficient stand-in for that object during pretraining
- domain assumption Automated captioning, entity extraction, grounding, and segmentation tools yield correspondences accurate enough at 773K scale for grounding supervision to help
- domain assumption Next-token prediction and entity reconstruction conditioned on visual patch tokens produce semantic grounding rather than shallow co-occurrence statistics
Cite this review
Pith. "Pith review of MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment." pith.science (2026). https://pith.science/paper/T5QJFT6E
@misc{pith2026260811167,
author = {Pith},
title = {Pith review of: MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment},
year = {2026},
howpublished = {\url{https://pith.science/paper/T5QJFT6E}},
note = {Machine review of arXiv:2608.11167}
}
read the original abstract
Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment suffers from referential ambiguity: models struggle to infer the correspondences between multiple visual objects and textual entities from the global representation, leading to data inefficiency and suboptimal semantic grounding. To address this, we propose MultiModal Code-Switching (MMCS), a novel pretraining paradigm that provides explicit object-level supervision. Inspired by the linguistic phenomenon of code-switching, MMCS interleaves vision and language by replacing textual entities with their corresponding visual objects, enforcing local vision-language grounding. We further develop a scalable data synthesis pipeline to generate a pretraining dataset of 773K samples with accurate object-entity correspondences. Experiments show that MMCS is highly data-efficient: with only 50K samples, it matches or surpasses models trained on 600K image-text pairs. Furthermore, MMCS consistently improves visual grounding and perception capabilities across varying model scales.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Visual Instruction Tuning , booktitle =
Haotian Liu and Chunyuan Li and Qingyang Wu and Yong Jae Lee , editor =. Visual Instruction Tuning , booktitle =. 2023 , url =
2023
-
[3]
LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , url=
Liu, Haotian and Li, Chunyuan and Li, Yuheng and Li, Bo and Zhang, Yuanhan and Shen, Sheng and Lee, Yong Jae , month=. LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , url=
-
[4]
Wenliang Dai and Junnan Li and Dongxu Li and Anthony Meng Huat Tiong and Junqi Zhao and Weisheng Wang and Boyang Li and Pascale Fung and Steven C. H. Hoi , editor =. InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning , booktitle =. 2023 , url =
2023
-
[5]
Peng Wang and Shuai Bai and Sinan Tan and Shijie Wang and Zhihao Fan and Jinze Bai and Keqin Chen and Xuejing Liu and Jialin Wang and Wenbin Ge and Yang Fan and Kai Dang and Mengfei Du and Xuancheng Ren and Rui Men and Dayiheng Liu and Chang Zhou and Jingren Zhou and Junyang Lin , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2409.12191 , epr...
-
[6]
Qwen2.5-VL Technical Report , journal =
Shuai Bai and Keqin Chen and Xuejing Liu and Jialin Wang and Wenbin Ge and Sibo Song and Kai Dang and Peng Wang and Shijie Wang and Jun Tang and Humen Zhong and Yuanzhi Zhu and Ming. Qwen2.5-VL Technical Report , journal =. 2025 , url =. doi:10.48550/ARXIV.2502.13923 , eprinttype =. 2502.13923 , timestamp =
-
[7]
2025 , eprint=
Qwen3-VL Technical Report , author=. 2025 , eprint=
2025
-
[8]
MiMo-VL Technical Report , journal =
Zihao Yue and Zhenru Lin and Yifan Song and Weikun Wang and Shuhuai Ren and Shuhao Gu and Shicheng Li and Peidian Li and Liang Zhao and Lei Li and Kainan Bao and Hao Tian and Hailin Zhang and Xiao. MiMo-VL Technical Report , journal =. 2025 , url =. doi:10.48550/ARXIV.2506.03569 , eprinttype =. 2506.03569 , timestamp =
-
[9]
2025 , eprint=
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning , author=. 2025 , eprint=
2025
Show all 87 references
-
[10]
CoRR , volume =
Zhe Chen and Weiyun Wang and Yue Cao and Yangzhou Liu and Zhangwei Gao and Erfei Cui and Jinguo Zhu and Shenglong Ye and Hao Tian and Zhaoyang Liu and Lixin Gu and Xuehui Wang and Qingyun Li and Yimin Ren and Zixuan Chen and Jiapeng Luo and Jiahao Wang and Tan Jiang and Bo Wan...
-
[11]
CoRR , volume =
Jinguo Zhu and Weiyun Wang and Zhe Chen and Zhaoyang Liu and Shenglong Ye and Lixin Gu and Hao Tian and Yuchen Duan and Weijie Su and Jie Shao and Zhangwei Gao and Erfei Cui and Xuehui Wang and Yue Cao and Yangzhou Liu and Xingguang Wei and Hongjie Zhang and Haomin Wang and We...
-
[12]
Scalable Vision Language Model Training via High Quality Data Curation , booktitle =
Hongyuan Dong and Zijian Kang and Weijie Yin and LiangXiao LiangXiao and ChaoFeng ChaoFeng and Ran Jiao , editor =. Scalable Vision Language Model Training via High Quality Data Curation , booktitle =. 2025 , url =
2025
-
[13]
Flamingo: a Visual Language Model for Few-Shot Learning , booktitle =
Jean. Flamingo: a Visual Language Model for Few-Shot Learning , booktitle =. 2022 , url =
2022
- [14]
-
[15]
CoRR , volume =
Dong Guo and Faming Wu and Feida Zhu and Fuxing Leng and Guang Shi and Haobin Chen and Haoqi Fan and Jian Wang and Jianyu Jiang and Jiawei Wang and Jingji Chen and Jingjia Huang and Kang Lei and Liping Yuan and Lishu Luo and Pengfei Liu and Qinghao Ye and Rui Qian and Shen Yan...
-
[16]
CoRR , volume =
Haiwen Diao and Mingxuan Li and Silei Wu and Linjun Dai and Xiaogang Wang and Hanming Deng and Lewei Lu and Dahua Lin and Ziwei Liu , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2510.14979 , eprinttype =. 2510.14979 , timestamp =
2025 doi
-
[17]
Computer Vision -
Brandon McKinzie and Zhe Gan and Jean. Computer Vision -. 2024 , url =. doi:10.1007/978-3-031-73397-0\_18 , timestamp =
2024 doi
-
[18]
CoRR , volume =
An Yang and Baosong Yang and Beichen Zhang and Binyuan Hui and Bo Zheng and Bowen Yu and Chengyuan Li and Dayiheng Liu and Fei Huang and Haoran Wei and Huan Lin and Jian Yang and Jianhong Tu and Jianwei Zhang and Jianxin Yang and Jiaxi Yang and Jingren Zhou and Junyang Lin and...
-
[19]
CoRR , volume =
An Yang and Anfeng Li and Baosong Yang and Beichen Zhang and Binyuan Hui and Bo Zheng and Bowen Yu and Chang Gao and Chengen Huang and Chenxu Lv and Chujie Zheng and Dayiheng Liu and Fan Zhou and Fei Huang and Feng Hu and Hao Ge and Haoran Wei and Huan Lin and Jialong Tang and...
- [20]
-
[21]
Making the
Yash Goyal and Tejas Khot and Douglas Summers. Making the. 2017. 2017 , url =. doi:10.1109/CVPR.2017.670 , timestamp =
2017 doi
-
[22]
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models , booktitle =
Kaichen Zhang and Bo Li and Peiyuan Zhang and Fanyi Pu and Joshua Adrian Cahyono and Kairui Hu and Shuai Liu and Yuanhan Zhang and Jingkang Yang and Chunyuan Li and Ziwei Liu , editor =. LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models , booktitle =. 2025 ...
2025 doi
-
[23]
MMBench: Is Your Multi-modal Model an All-Around Player? , booktitle =
Yuan Liu and Haodong Duan and Yuanhan Zhang and Bo Li and Songyang Zhang and Wangbo Zhao and Yike Yuan and Jiaqi Wang and Conghui He and Ziwei Liu and Kai Chen and Dahua Lin , editor =. MMBench: Is Your Multi-modal Model an All-Around Player? , booktitle =. 2024 , url =. doi:1...
2024 doi
-
[24]
CoRR , volume =
Chaoyou Fu and Peixian Chen and Yunhang Shen and Yulei Qin and Mengdan Zhang and Xu Lin and Zhenyu Qiu and Wei Lin and Jinrui Yang and Xiawu Zheng and Ke Li and Xing Sun and Rongrong Ji , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2306.13394 , eprinttype =. 230...
-
[25]
Are We on the Right Way for Evaluating Large Vision-Language Models? , booktitle =
Lin Chen and Jinsong Li and Xiaoyi Dong and Pan Zhang and Yuhang Zang and Zehui Chen and Haodong Duan and Jiaqi Wang and Yu Qiao and Dahua Lin and Feng Zhao , editor =. Are We on the Right Way for Evaluating Large Vision-Language Models? , booktitle =. 2024 , url =
2024
-
[27]
Forty-first International Conference on Machine Learning,
Weihao Yu and Zhengyuan Yang and Linjie Li and Jianfeng Wang and Kevin Lin and Zicheng Liu and Xinchao Wang and Lijuan Wang , title =. Forty-first International Conference on Machine Learning,. 2024 , url =
2024
-
[28]
Hudson and Christopher D
Drew A. Hudson and Christopher D. Manning , title =. 2019 , url =. doi:10.1109/CVPR.2019.00686 , timestamp =
2019
-
[29]
A Diagram is Worth a Dozen Images , booktitle =
Aniruddha Kembhavi and Mike Salvato and Eric Kolve and Min Joon Seo and Hannaneh Hajishirzi and Ali Farhadi , editor =. A Diagram is Worth a Dozen Images , booktitle =. 2016 , url =. doi:10.1007/978-3-319-46493-0\_15 , timestamp =
2016 doi
-
[30]
Joty and Enamul Hoque , editor =
Ahmed Masry and Do Xuan Long and Jia Qing Tan and Shafiq R. Joty and Enamul Hoque , editor =. ChartQA:. Findings of the Association for Computational Linguistics:. 2022 , url =. doi:10.18653/V1/2022.FINDINGS-ACL.177 , timestamp =
2022 doi
-
[31]
Cambrian-1:
Peter Tong and Ellis Brown and Penghao Wu and Sanghyun Woo and Adithya Iyer and Sai Charitha Akula and Shusheng Yang and Jihan Yang and Manoj Middepogu and Ziteng Wang and Xichen Pan and Rob Fergus and Yann LeCun and Saining Xie , editor =. Cambrian-1:. Advances in Neural Info...
2024
-
[32]
OCRBench: on the hidden mystery of
Yuliang Liu and Zhang Li and Mingxin Huang and Biao Yang and Wenwen Yu and Chunyuan Li and Xu. OCRBench: on the hidden mystery of. Sci. China Inf. Sci. , volume =. 2024 , url =. doi:10.1007/S11432-024-4235-6 , timestamp =
2024 doi
-
[33]
2019 , url =
Amanpreet Singh and Vivek Natarajan and Meet Shah and Yu Jiang and Xinlei Chen and Dhruv Batra and Devi Parikh and Marcus Rohrbach , title =. 2019 , url =. doi:10.1109/CVPR.2019.00851 , timestamp =
2019
-
[35]
Berg , editor =
Sahar Kazemzadeh and Vicente Ordonez and Mark Matten and Tamara L. Berg , editor =. ReferItGame: Referring to Objects in Photographs of Natural Scenes , booktitle =. 2014 , url =. doi:10.3115/V1/D14-1086 , timestamp =
2014 doi
-
[36]
Yuille and Kevin Murphy , title =
Junhua Mao and Jonathan Huang and Alexander Toshev and Oana Camburu and Alan L. Yuille and Kevin Murphy , title =. 2016. 2016 , url =. doi:10.1109/CVPR.2016.9 , timestamp =
2016 doi
-
[37]
Minesh Mathew and Dimosthenis Karatzas and C. V. Jawahar , title =. 2021 , url =. doi:10.1109/WACV48630.2021.00225 , timestamp =
2021
-
[38]
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models , booktitle =
Matt Deitke and Christopher Clark and Sangho Lee and Rohun Tripathi and Yue Yang and Jae Sung Park and Mohammadreza Salehi and Niklas Muennighoff and Kyle Lo and Luca Soldaini and Jiasen Lu and Taira Anderson and Erin Bransom and Kiana Ehsani and Huong Ngo and Yen. Molmo and P...
2025
-
[39]
ShareGPT4V: Improving Large Multi-modal Models with Better Captions , booktitle =
Lin Chen and Jinsong Li and Xiaoyi Dong and Pan Zhang and Conghui He and Jiaqi Wang and Feng Zhao and Dahua Lin , editor =. ShareGPT4V: Improving Large Multi-modal Models with Better Captions , booktitle =. 2024 , url =. doi:10.1007/978-3-031-72643-9\_22 , timestamp =
2024 doi
-
[40]
CoRR , volume =
Guiming Hardy Chen and Shunian Chen and Ruifei Zhang and Junying Chen and Xiangbo Wu and Zhiyi Zhang and Zhihong Chen and Jianquan Li and Xiang Wan and Benyou Wang , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2402.11684 , eprinttype =. 2402.11684 , timestamp =
-
[41]
DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception , booktitle =
Xiaotong Li and Fan Zhang and Haiwen Diao and Yueze Wang and Xinlong Wang and Lingyu Duan , editor =. DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception , booktitle =. 2024 , url =
2024
-
[42]
CoRR , volume =
Xiangtai Li and Tao Zhang and Yanwei Li and Haobo Yuan and Shihao Chen and Yikang Zhou and Jiahao Meng and Yueyi Sun and Shilin Xu and Lu Qi and Tianheng Cheng and Yi Lin and Zilong Huang and Wenhao Huang and Jiashi Feng and Guang Shi , title =. CoRR , volume =. 2025 , url =. ...
-
[43]
ImageInWords: Unlocking Hyper-Detailed Image Descriptions , booktitle =
Roopal Garg and Andrea Burns and Burcu Karagol Ayan and Yonatan Bitton and Ceslee Montgomery and Yasumasa Onoe and Andrew Bunner and Ranjay Krishna and Jason Baldridge and Radu Soricut , editor =. ImageInWords: Unlocking Hyper-Detailed Image Descriptions , booktitle =. 2024 , ...
2024 doi
-
[44]
Computer Vision -
Yasumasa Onoe and Sunayana Rane and Zachary Berger and Yonatan Bitton and Jaemin Cho and Roopal Garg and Alexander Ku and Zarana Parekh and Jordi Pont. Computer Vision -. 2024 , url =. doi:10.1007/978-3-031-73027-6\_17 , timestamp =
2024 doi
-
[45]
CoRR , volume =
Long Xing and Xiaoyi Dong and Yuhang Zang and Yuhang Cao and Jianze Liang and Qidong Huang and Jiaqi Wang and Feng Wu and Dahua Lin , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2509.22647 , eprinttype =. 2509.22647 , timestamp =
2025 doi
-
[46]
Microsoft
Tsung. Microsoft. Computer Vision -. 2014 , url =. doi:10.1007/978-3-319-10602-1\_48 , timestamp =
2014 doi
-
[47]
Shuai Shao and Zeming Li and Tianyuan Zhang and Chao Peng and Gang Yu and Xiangyu Zhang and Jing Li and Jian Sun , title =. 2019. 2019 , url =. doi:10.1109/ICCV.2019.00852 , timestamp =
2019
-
[48]
Alina Kuznetsova and Hassan Rom and Neil Alldrin and Jasper R. R. Uijlings and Ivan Krasin and Jordi Pont. The Open Images Dataset. Int. J. Comput. Vis. , volume =. 2020 , url =. doi:10.1007/S11263-020-01316-Z , timestamp =
2020 doi
-
[49]
2024 , url =
Youcai Zhang and Xinyu Huang and Jinyu Ma and Zhaoyang Li and Zhaochuan Luo and Yanchun Xie and Yuzhuo Qin and Tong Luo and Yaqian Li and Shilong Liu and Yandong Guo and Lei Zhang , title =. 2024 , url =. doi:10.1109/CVPRW63382.2024.00179 , timestamp =
2024
-
[50]
Segment Anything , booktitle =
Alexander Kirillov and Eric Mintun and Nikhila Ravi and Hanzi Mao and Chlo. Segment Anything , booktitle =. 2023 , url =. doi:10.1109/ICCV51070.2023.00371 , timestamp =
2023
-
[51]
Peter Young and Alice Lai and Micah Hodosh and Julia Hockenmaier , title =. Trans. Assoc. Comput. Linguistics , volume =. 2014 , url =. doi:10.1162/TACL\_A\_00166 , timestamp =
2014 doi
-
[52]
2019 International Conference on Document Analysis and Recognition,
Anand Mishra and Shashank Shekhar and Ajeet Kumar Singh and Anirban Chakraborty , title =. 2019 International Conference on Document Analysis and Recognition,. 2019 , url =. doi:10.1109/ICDAR.2019.00156 , timestamp =
2019
-
[53]
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations , journal =
Ranjay Krishna and Yuke Zhu and Oliver Groth and Justin Johnson and Kenji Hata and Joshua Kravitz and Stephanie Chen and Yannis Kalantidis and Li. Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations , journal =. 2017 , url =. doi:10.1007/S1...
2017 doi
-
[54]
Price and Scott Cohen and Christopher Kanan , title =
Kushal Kafle and Brian L. Price and Scott Cohen and Christopher Kanan , title =. 2018. 2018 , url =. doi:10.1109/CVPR.2018.00592 , timestamp =
2018
-
[55]
OCR-Free Document Understanding Transformer , booktitle =
Geewook Kim and Teakgyu Hong and Moonbin Yim and JeongYeon Nam and Jinyoung Park and Jinyeong Yim and Wonseok Hwang and Sangdoo Yun and Dongyoon Han and Seunghyun Park , editor =. OCR-Free Document Understanding Transformer , booktitle =. 2022 , url =. doi:10.1007/978-3-031-19...
2022 doi
-
[56]
Christoph Schuhmann and Romain Beaumont and Richard Vencu and Cade Gordon and Ross Wightman and Mehdi Cherti and Theo Coombes and Aarush Katta and Clayton Mullis and Mitchell Wortsman and Patrick Schramowski and Srivatsa Kundurthy and Katherine Crowson and Ludwig Schmidt and R...
2022
-
[57]
Grounding
Shilong Liu and Zhaoyang Zeng and Tianhe Ren and Feng Li and Hao Zhang and Jie Yang and Qing Jiang and Chunyuan Li and Jianwei Yang and Hang Su and Jun Zhu and Lei Zhang , editor =. Grounding. Computer Vision -. 2024 , url =. doi:10.1007/978-3-031-72970-6\_3 , timestamp =
2024 doi
-
[58]
The Thirteenth International Conference on Learning Representations,
Nikhila Ravi and Valentin Gabeur and Yuan. The Thirteenth International Conference on Learning Representations,. 2025 , url =
2025
-
[59]
Forty-first International Conference on Machine Learning,
Minyoung Huh and Brian Cheung and Tongzhou Wang and Phillip Isola , title =. Forty-first International Conference on Machine Learning,. 2024 , url =
2024
-
[60]
Hinton , editor =
Simon Kornblith and Mohammad Norouzi and Honglak Lee and Geoffrey E. Hinton , editor =. Similarity of Neural Network Representations Revisited , booktitle =. 2019 , url =
2019
-
[61]
SEA : Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLM s
Yin, Yuanyang and Zhao, Yaqi and Zhang, Yajie and Zhang, Yuanxing and Lin, Ke and Wang, Jiahao and Tao, Xin and Wan, Pengfei and Zhang, Wentao and Zhao, Feng. SEA : Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLM s. Proceedings of the 2025 Con...
2025 doi
- [62]
-
[64]
GroundingGPT: Language Enhanced Multi-modal Grounding Model , booktitle =
Zhaowei Li and Qi Xu and Dong Zhang and Hang Song and Yiqing Cai and Qi Qi and Ran Zhou and Junting Pan and Zefeng Li and Vu Tu and Zhida Huang and Tao Wang , editor =. GroundingGPT: Language Enhanced Multi-modal Grounding Model , booktitle =. 2024 , url =. doi:10.18653/V1/202...
2024 doi
-
[65]
Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models
Wang, Wei and Li, Zhaowei and Xu, Qi and Li, Linfeng and Cai, YiQing and Jiang, Botian and Song, Hang and Hu, Xingcan and Wang, Pengyu and Xiao, Li. Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models. Proceedings of the 2025 Conference...
2025 doi
- [66]
-
[68]
ParGo: Bridging Vision-Language with Partial and Global Views , booktitle =
An. ParGo: Bridging Vision-Language with Partial and Global Views , booktitle =. 2025 , url =. doi:10.1609/AAAI.V39I7.32806 , timestamp =
2025 doi
- [69]
-
[70]
Ferret: Refer and Ground Anything Anywhere at Any Granularity , booktitle =
Haoxuan You and Haotian Zhang and Zhe Gan and Xianzhi Du and Bowen Zhang and Zirui Wang and Liangliang Cao and Shih. Ferret: Refer and Ground Anything Anywhere at Any Granularity , booktitle =. 2024 , url =
2024
-
[71]
ICCV , year=
Region-based Cluster Discrimination for Visual Representation Learning , author=. ICCV , year=
-
[72]
Bo Li and Yuanhan Zhang and Dong Guo and Renrui Zhang and Feng Li and Hao Zhang and Kaichen Zhang and Peiyuan Zhang and Yanwei Li and Ziwei Liu and Chunyuan Li , title =. Trans. Mach. Learn. Res. , volume =. 2025 , url =
2025
-
[73]
CoRR , volume =
Xiang An and Yin Xie and Kaicheng Yang and Wenkang Zhang and Xiuwei Zhao and Zheng Cheng and Yirui Wang and Songcen Xu and Changrui Chen and Chunsheng Wu and Huajie Tan and Chunyuan Li and Jing Yang and Jie Yu and Xiyao Wang and Bin Qin and Yumeng Wang and Zizhen Yan and Ziyon...
-
[74]
Learning Transferable Visual Models From Natural Language Supervision , booktitle =
Alec Radford and Jong Wook Kim and Chris Hallacy and Aditya Ramesh and Gabriel Goh and Sandhini Agarwal and Girish Sastry and Amanda Askell and Pamela Mishkin and Jack Clark and Gretchen Krueger and Ilya Sutskever , editor =. Learning Transferable Visual Models From Natural La...
2021
-
[75]
Gritsenko and Xiao Wang and Muhammad Ferjad Naeem and Ibrahim Alabdulmohsin and Nikhil Parthasarathy and Talfan Evans and Lucas Beyer and Ye Xia and Basil Mustafa and Olivier J
Michael Tschannen and Alexey A. Gritsenko and Xiao Wang and Muhammad Ferjad Naeem and Ibrahim Alabdulmohsin and Nikhil Parthasarathy and Talfan Evans and Lucas Beyer and Ye Xia and Basil Mustafa and Olivier J. H. SigLIP 2: Multilingual Vision-Language Encoders with Improved Se...
-
[76]
GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest , booktitle =
Shilong Zhang and Peize Sun and Shoufa Chen and Min Xiao and Wenqi Shao and Wenwei Zhang and Yu Liu and Kai Chen and Ping Luo , editor =. GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest , booktitle =. 2024 , url =. doi:10.1007/978-3-031-91813-1\_4 , timestamp =
2024 doi
-
[77]
Shaker and Salman H
Hanoona Abdul Rasheed and Muhammad Maaz and Sahal Shaji Mullappilly and Abdelrahman M. Shaker and Salman H. Khan and Hisham Cholakkal and Rao Muhammad Anwer and Eric P. Xing and Ming. GLaMM: Pixel Grounding Large Multimodal Model , booktitle =. 2024 , url =. doi:10.1109/CVPR52...
2024
-
[78]
NExT-Chat: An
Ao Zhang and Yuan Yao and Wei Ji and Zhiyuan Liu and Tat. NExT-Chat: An. Forty-first International Conference on Machine Learning,. 2024 , url =
2024
-
[79]
LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models , booktitle =
Hao Zhang and Hongyang Li and Feng Li and Tianhe Ren and Xueyan Zou and Shilong Liu and Shijia Huang and Jianfeng Gao and Leizhang and Chunyuan Li and Jainwei Yang , editor =. LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models , booktitle =. 2024 , url =. doi:1...
2024 doi
-
[80]
The All-Seeing Project
Weiyun Wang and Yiming Ren and Haowen Luo and Tiantong Li and Chenxiang Yan and Zhe Chen and Wenhai Wang and Qingyun Li and Lewei Lu and Xizhou Zhu and Yu Qiao and Jifeng Dai , editor =. The All-Seeing Project. Computer Vision -. 2024 , url =. doi:10.1007/978-3-031-73414-4\_27...
2024 doi
-
[81]
Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training , booktitle =
Zhijun Wang and Jiahuan Li and Hao Zhou and Rongxiang Weng and Jingang Wang and Xin Huang and Xue Han and Junlan Feng and Chao Deng and Shujian Huang , editor =. Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training , booktitle =. 2025 , url =
2025
-
[82]
PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual Alignment , booktitle =
Jiahuan Li and Shujian Huang and Aarron Ching and Xinyu Dai and Jiajun Chen , editor =. PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual Alignment , booktitle =. 2024 , url =. doi:10.18653/V1/2024.EMNLP-MAIN.572 , timestamp =
2024 doi
-
[83]
Code-Switching Curriculum Learning for Multilingual Transfer in LLMs , booktitle =
Haneul Yoo and Cheonbok Park and Sangdoo Yun and Alice Oh and Hwaran Lee , editor =. Code-Switching Curriculum Learning for Multilingual Transfer in LLMs , booktitle =. 2025 , url =
2025
-
[84]
CoRR , volume =
Aditi Chaudhary and Karthik Raman and Krishna Srinivasan and Jiecao Chen , title =. CoRR , volume =. 2020 , url =. 2010.12566 , timestamp =
2020 arXiv
-
[85]
CoSDA-ML: Multi-Lingual Code-Switching Data Augmentation for Zero-Shot Cross-Lingual
Libo Qin and Minheng Ni and Yue Zhang and Wanxiang Che , editor =. CoSDA-ML: Multi-Lingual Code-Switching Data Augmentation for Zero-Shot Cross-Lingual. Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence,. 2020 , url =. doi:10.24963/IJCAI...
2020 doi
-
[86]
Bilingual Alignment Pre-Training for Zero-Shot Cross-Lingual Transfer
Yang, Ziqing and Ma, Wentao and Cui, Yiming and Ye, Jiani and Che, Wanxiang and Wang, Shijin. Bilingual Alignment Pre-Training for Zero-Shot Cross-Lingual Transfer. Proceedings of the 3rd Workshop on Machine Reading for Question Answering. 2021. doi:10.18653/v1/2021.mrqa-1.10
2021 doi
-
[87]
Latino Language and Communicative Behavior , editor =
Syntactic structure and social function of code-switching , author =. Latino Language and Communicative Behavior , editor =. 1981 , publisher =
1981
-
[88]
Thara and Prabaharan Poornachandran , title =
S. Thara and Prabaharan Poornachandran , title =. 2018 International Conference on Advances in Computing, Communications and Informatics,. 2018 , url =. doi:10.1109/ICACCI.2018.8554413 , timestamp =
2018
-
[89]
Alabdulmohsin and Avital Oliver and Piotr Padlewski and Alexey A
Mostafa Dehghani and Basil Mustafa and Josip Djolonga and Jonathan Heek and Matthias Minderer and Mathilde Caron and Andreas Steiner and Joan Puigcerver and Robert Geirhos and Ibrahim M. Alabdulmohsin and Avital Oliver and Piotr Padlewski and Alexey A. Gritsenko and Mario Luci...
2023
-
[90]
Hu and Yelong Shen and Phillip Wallis and Zeyuan Allen
Edward J. Hu and Yelong Shen and Phillip Wallis and Zeyuan Allen. LoRA: Low-Rank Adaptation of Large Language Models , booktitle =. 2022 , url =
2022
-
[91]
2025 , howpublished =
Google DeepMind , title =. 2025 , howpublished =
2025
-
[92]
18th International Conference on Pattern Recognition
Alexander Neubeck and Luc Van Gool , title =. 18th International Conference on Pattern Recognition. 2006 , url =. doi:10.1109/ICPR.2006.479 , timestamp =
2006 doi
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.