REVIEW 4 major objections 4 minor 113 references
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read POBF repaints only the background around a target object, then filters synthetic samples with teacher scores, yielding an average 5.83% gain over real-data-only visual grounding training in data-scarce settings.
desk verdict A useful synthetic-data recipe for data-scarce visual grounding, with a plausible label-alignment story that needs verification. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the paint-outside-the-box inpainting step: an off-the-shelf diffusion inpainting model regenerates everything outside the bounding box at high strength while the pixels inside the box are copied unchanged, so the synthetic image is guaranteed to align with the box by construction. On top of that sits a three-term selection score computed by a teacher network trained only on the scarce real data. The hardness score $S_1$ is the teacher's IoU with the true box given the text; the overfitting score $S_2 = 1 - \mathrm{IoU}(\mathrm{teacher}(\mathrm{masked\ image}, \mathrm{text}), \mathrm{box})$ flags backgrounds that leak the answer; and the penalty term $P$ is the IoU the teacher gets from the image alone with an empty text prompt. Together they rank the generated variants and one winner per real sample enters the student's training set.
What would settle it
Compare POBF's selected synthetic set against a version in which each generated image's box content is validated by an independent object detector or human annotator; if removing samples with corrupted box contents does not lower accuracy or the gain disappears, the assumed label alignment is not what drives the result.
Extended reading notes
Core claim
The central claim is that the label misalignment that plagues synthetic visual grounding data comes from editing the object rather than the context, and that painting outside the box fixes it. Starting from a real image, a real text query, and its bounding box, the method captions the image, inpaints the region outside the box with high strength so the background changes substantially, and pairs the new image with the original box and text. Because the box pixels are never regenerated, the synthetic sample inherits the exact ground-truth localization. A teacher trained on the scarce real data then scores each generated image: $S_1$ measures how confidently the teacher localizes the target with the text, $S_2$ measures whether the teacher can still localize it when the box is masked (a sign the background carries unintended shortcut features), and $P$ measures localization from the image alone with an empty text prompt. The weighted sum of the three normalized scores, tuned by grid search, picks one synthetic image per real sample, and the student trains on real plus selected synthetic data, sometimes with regenerated captions.
Load-bearing premise
The approach assumes that inpainting the background leaves the object inside the bounding box visually unchanged and semantically aligned with the original text; if the generator alters the object or introduces overlapping artifacts, the synthetic labels are wrong even after filtering.
Editorial extensions
If this is right
- If the claim holds, dense region-text annotations are not a hard requirement for visual grounding: a small real set plus repainted backgrounds and teacher-filtered selection can lift accuracy by 5.83% on average.
- Because the box pixels are never regenerated, label misalignment is avoided by construction, which is the paper's diagnosis of why prior object-editing and text-to-image baselines underperform.
- The filter's three scores beat the common CLIP-similarity selection rule by an average of 1.32%, suggesting that hardness and overfitting capture complementary quality signals.
- The gains are not tied to one generator or architecture: the paper reports consistent improvements across two alternative image generators, two alternative captioners, three grounding models, and three data-scarce budgets at 0.5%, 1%, and 2% of real data.
Reading between the lines
- The same paint-outside-the-box rule should transfer to other region-labeled tasks such as instance segmentation, detection, and keypoint localization: preserve the annotated region and regenerate its context, then filter by a teacher; this extension is not tested in the paper.
- The per-sample 'keep one winner' rule limits synthetic expansion to 2x; a diversity-aware selection that keeps multiple high-scoring variants could use the same scores to get more benefit from a fixed generator budget.
- Because the teacher is trained on the same tiny real set, the filter inherits the teacher's blind spots; averaging scores across multiple teacher seeds or architectures could make the selection more stable, but the paper's teacher set is single.
- The paper's own hint that small objects gain less suggests a testable extension: constrain the inpainted background area or use an object-aware mask so the generator does not have to synthesize huge surrounding regions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes POBF, a framework for data-scarce visual grounding. POBF generates synthetic training images by inpainting the region outside the ground-truth bounding box while keeping the boxed object intact, and it augments text by captioning the cropped box. A teacher model trained only on the limited real data then scores each synthetic image with a hardness score (IoU of the teacher's prediction with the box), an overfitting score (1 minus IoU when the box is masked), and a penalty term (IoU with an empty text query); the highest-scoring synthetic image per real sample is kept and added to the training set. Experiments on RefCOCO, RefCOCO+, RefCOCOg, and ReferItGame report an average gain of 5.83% over real-data-only training and consistent gains across different generators, captioners, data sizes, and student architectures.
Significance. If the empirical claims hold, POBF is a practical and flexible data-augmentation method for visual grounding under severe data scarcity. Its strengths are that the generation pipeline uses only off-the-shelf image-caption pretrained models, the filtering scheme is compared against several data-selection baselines, and the robustness experiments span four datasets, three architectures, three generators/captioners, and three training-data sizes. The paper is also transparent about the teacher being trained on real data only and about the held-out evaluation protocol. The main unresolved risk is that the central label-alignment assumption is asserted rather than verified, and the ablation table as printed does not support all of the stated component-wise gains.
major comments (4)
- [§4.1 and §5 Implementation Details] Label alignment is asserted rather than verified. The claim that painting outside the box leaves the object 'unchanged' and therefore 'strictly align[s] with the ground truth bounding box' is not demonstrated for the chosen strength of 0.9, which adds substantial noise to the base image and can alter pixels inside the box. Even if the box pixels were exactly preserved, replacing the background can invalidate relational referring expressions such as 'right cow' or 'bus in front of a building,' so the original text T may no longer uniquely refer to B in the new image. The hardness score in Eq. (1) checks only whether the teacher predicts B from T; it does not verify that T is a valid description of the boxed object in the synthetic image. I request a label-validity check (human evaluation on a sample, or an automated VQA/CLIP-based check) and a report of how often the box content changes at strength=0.9, because the 5.83% gain is attributed to a generation strategy whose labels are assumed correct.
- [Table 2] The headline comparison against X-Paste, Gen2Det, and GeoDiffusion is based on a single randomly sampled subset with no error bars, while the three-subset average is reported only for the Real and POBF rows. Thus the claim of outperforming leading baselines by 2.29%–3.85% rests on one draw. Please report the three-subset average (or at least standard deviations) for the baselines as well, or provide a significance test. In addition, the header 'All methods employ the proposed filtering scheme' needs a clear statement of which hyperparameters (K, q, lambda) were shared by the baselines, since the table otherwise reads as a comparison of full pipelines rather than generation strategies under a common filter.
- [§5.2, Table 3] The ablation table does not support the stated per-component gains. The text claims that S1, S2, and P give gains of 1.23%, 0.71%, and 1.14% 'compared to the variant without any filtering,' but Table 3 contains no visible row for S2 alone or P alone, and the differences between the no-filter Real+Synimg row (31.31) and the filtering rows are not 1.23, 0.71, and 1.14. The reported 'improvement of 1.96% over the baseline without filtering' for the full scheme also does not match any visible no-filter baseline; the only visible pairwise difference equal to 1.96 is with the S1-only row, which is itself a filtering variant. Please clarify the row encodings or add the missing rows so that each component's individual contribution can be verified.
- [§4.2, Eq. (3)] The penalty term P is computed as IoU(T(I', ∅), B), but TransVG and most grounding models require a non-empty text query. The manuscript does not specify how the empty string is encoded or why a box prediction with no text measures 'prior knowledge.' Since the penalty term is one of the three components of the final score in Eq. (4), this operational detail directly affects reproducibility of the filtering scheme.
minor comments (4)
- [Abstract vs. Conclusion] The reported margins over baselines are inconsistent: the Abstract and §5.1 state 2.29%–3.85%, while the Conclusion states 2.74%–4.35%; please correct one of these.
- [§4.2, Eq. (4)] Please specify the population over which the three scores are normalized; if the normalization is per real image (over its K synthetic variants) rather than global, the selection behavior of Eq. (4) is different and should be described.
- [§5.1] The explanation that the modest improvement on ReferIt stems from small objects and large inpainted background regions is a hypothesis; either provide supporting evidence (for example, an object-size analysis) or clearly mark it as speculative.
- [Table 5] The rows labeled 'Replace the Image Captioner with ...' do not show the default BLIP captioner row in the table; adding it would make the comparison easier to read.
Circularity Check
No significant circularity; the central gains are empirical comparisons against held-out benchmarks with independently defined filtering scores.
full rationale
POBF's derivation chain is not circular. The generation strategy (paint outside the box) is a data-augmentation recipe, not a theorem: the claim that the object inside the box remains aligned is an empirical assumption about the inpainting model, and any failure of that assumption is a correctness risk rather than a logical circularity. The filtering scores S1, S2, and P are defined through a teacher model trained only on the limited real data, and the final student is evaluated on held-out test splits, so the reported 5.83% gain over the real-only baseline is an external empirical result rather than a quantity forced by the definitions. The only fitted parameters are the three lambda weights, which are tuned on the validation set, not on the test set, and the scores themselves are not constructed from the student's test performance. No equation in the paper reduces to its own input, no self-citation carries the central load, and no known result is merely renamed. The paper's own limitation regarding label alignment at high inpainting strength is a concern about validity of synthetic labels, not circularity.
Assumptions & free parameters
free parameters (5)
- lambda_1 (hardness score weight) =
1.0 or 0.5 via grid search on validation; exact value per dataset not reported
- lambda_2 (overfitting score weight) =
1.0 or 0.5 via grid search on validation; exact value per dataset not reported
- lambda_P (penalty term weight) =
1.0 or 0.5 via grid search on validation; exact value per dataset not reported
- K (synthetic images per real sample) =
4
- q (caption replacement probability) =
0.3
assumptions (4)
- domain assumption The inpainting model preserves the content inside the bounding box exactly when painting outside, so the original box remains a valid label for the generated image.
- domain assumption A teacher model trained on 1% real data produces hardness and overfitting scores that rank synthetic samples by their usefulness for student training.
- domain assumption Selecting easier synthetic samples and samples that do not allow box prediction from background improves student generalization.
- domain assumption IoU-based top-1 accuracy with a 0.5 threshold is an appropriate proxy for visual grounding quality.
Cite this review
Pith. "Pith review of Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding." pith.science (2026). https://pith.science/paper/I3QVJPTU
@misc{pith2026241200684,
author = {Pith},
title = {Pith review of: Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding},
year = {2026},
howpublished = {\url{https://pith.science/paper/I3QVJPTU}},
note = {Machine review of arXiv:2412.00684}
}
read the original abstract
Visual grounding aims to localize the image regions based on a textual query. Given the difficulty of large-scale data curation, we investigate how to effectively learn visual grounding under data-scarce settings in this paper. To address the data scarcity, we propose a novel framework, POBF (Paint Outside the Box and Filter). POBF synthesizes images by inpainting outside the box, tackling a label misalignment issue encountered in previous works. Furthermore, POBF leverages an innovative filtering scheme to select the most effective training data. This scheme combines a hardness score and an overfitting score, balanced by a penalty term. Extensive experiments across four benchmark datasets demonstrate that POBF consistently improves performance, achieving an average gain of 5.83\% over the real-data-only method and outperforming leading baselines by 2.29\%-3.85\% in accuracy. Additionally, we validate the robustness and generalizability of POBF across various generative models, training data sizes, and model architectures.
Figures
Reference graph
Works this paper leans on
-
[1]
Amro Abbas, Kushal Tirumala, Dániel Simig, Surya Ganguli, and Ari S Mor- cos. 2023. Semdedup: Data-efficient learning at web-scale through semantic deduplication. arXiv preprint arXiv:2303.09540 (2023)
arXiv 2023
-
[2]
Sara Beery, Grant Van Horn, and Pietro Perona. 2018. Recognition in terra incognita. In Proceedings of the ECCV . 456–473
2018
-
[3]
Fangyi Chen, Han Zhang, Zhantao Yang, Hao Chen, Kai Hu, and Marios Sav- vides. 2024. RTGen: Generating Region-Text Pairs for Open-Vocabulary Object Detection. arXiv preprint arXiv:2405.19854 (2024)
work page Pith review arXiv 2024
-
[4]
Gongwei Chen, Leyang Shen, Rui Shao, Xiang Deng, and Liqiang Nie. 2024. Lion: Empowering multimodal large language model with dual-level visual knowledge. In Proceedings of the IEEE/CVF CVPR . 26540–26550
2024
-
[5]
Kai Chen, Enze Xie, Zhe Chen, Yibo Wang, HONG Lanqing, Zhenguo Li, and Dit-Yan Yeung. 2024. GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation. In The Twelfth International Conference on Learning Representations
2024
-
[6]
Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao
-
[7]
Pengfei Chen, Ben Ben Liao, Guangyong Chen, and Shengyu Zhang. 2019. Un- derstanding and utilizing deep neural networks trained with noisy labels. In International Conference on Machine Learning . PMLR, 1062–1070
2019
-
[8]
Sijia Chen and Baochun Li. 2022. Multi-modal dynamic graph transformer for visual grounding. In Proceedings of the IEEE/CVF CVPR . 15534–15543
2022
Show all 113 references
-
[9]
Sijia Chen and Baochun Li. 2023. Language-guided diffusion model for visual grounding. arXiv preprint arXiv:2308.09599 (2023)
2023 arXiv
-
[10]
Shengxin Chen, Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Guannan Jiang, and Ron- grong Ji. 2024. QueryMatch: A Query-based Contrastive Learning Framework for Weakly Supervised Visual Grounding. In Proceedings of the 32nd ACM Inter- national Conference on Multimedia . 4177–4186
2024
-
[11]
Yicheng Chen, Xiangtai Li, Yining Li, Yanhong Zeng, Jianzong Wu, Xiangyu Zhao, and Kai Chen. 2024. Auto Cherry-Picker: Learning from High-quality Generative Data Driven by Language. arXiv preprint arXiv:2406.20085 (2024)
2024 arXiv
-
[12]
Jiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou, and Houqiang Li
-
[13]
Jiajun Deng, Zhengyuan Yang, Daqing Liu, Tianlang Chen, Wengang Zhou, Yanyong Zhang, Houqiang Li, and Wanli Ouyang. 2023. Transvg++: End-to-end visual grounding with language conditioned vision transformer.IEEE transactions on pattern analysis and machine intelligence (2023)
2023
-
[14]
Zilin Du, Yunxin Li, Xu Guo, Yidan Sun, and Boyang Li. 2023. Training Multi- media Event Extraction With Generated Images and Captions. arXiv preprint arXiv:2306.08966 (2023)
2023 arXiv
-
[15]
Chengxiang Fan, Muzhi Zhu, Hao Chen, Yang Liu, Weijia Wu, Huaqi Zhang, and Chunhua Shen. 2024. DiverGen: Improving Instance Segmentation by Learning Wider Data Distribution with More Diverse Generative Data. In Proceedings of the IEEE/CVF CVPR. 3986–3995
2024
-
[16]
Haoyang Fang, Boran Han, Shuai Zhang, Su Zhou, Cuixiong Hu, and Wen-Ming Ye. 2024. Data augmentation for object detection via controllable diffusion models. In Proceedings of the IEEE/CVF W ACV. 1257–1266
2024
-
[17]
Chengjian Feng, Yujie Zhong, Zequn Jie, Weidi Xie, and Lin Ma. 2024. Instagen: Enhancing object detection by training on synthetic dataset. In Proceedings of the IEEE/CVF CVPR. 14121–14130
2024
-
[18]
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. 2020. Shortcut learning in deep neural networks. Nature Machine Intelligence 2, 11 (2020)
2020
-
[19]
Zeyu Han, Fangrui Zhu, Qianru Lao, and Huaizu Jiang. 2024. Zero-shot referring expression comprehension via structural similarity between images and captions. In Proceedings of the IEEE/CVF CVPR . 14364–14374
2024
-
[20]
Ruozhen He, Paola Cascante-Bonilla, Ziyan Yang, Alexander C Berg, and Vicente Ordonez. 2024. Improved Visual Grounding through Self-Consistent Explanations. In Proceedings of the IEEE/CVF CVPR . 13095–13105
2024
-
[21]
Ruozhen He, Paola Cascante-Bonilla, Ziyan Yang, Alexander C Berg, and Vicente Ordonez. 2024. Learning from Models and Data for Visual Grounding. arXiv preprint arXiv:2403.13804 (2024)
2024 arXiv
-
[22]
Haojun Jiang, Yuanze Lin, Dongchen Han, Shiji Song, and Gao Huang. 2022. Pseudo-q: Generating pseudo language queries for visual grounding. In Proceed- ings of the IEEE/CVF CVPR
2022
-
[23]
Lu Jiang, Zhengyuan Zhou, Thomas Leung, Li-Jia Li, and Li Fei-Fei. 2018. Mentor- net: Learning data-driven curriculum for very deep neural networks on corrupted labels. In International conference on machine learning . PMLR
2018
-
[24]
Yiding Jiang, Allan Zhou, Zhili Feng, Sadhika Malladi, and J Zico Kolter. 2024. Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws.arXiv preprint arXiv:2410.11820 (2024)
2024 arXiv
-
[25]
Yang Jiao, Zequn Jie, Jingjing Chen, Lin Ma, and Yu-Gang Jiang. 2023. Suspected Objects Matter: Rethinking Model’s Prediction for One-stage Visual Grounding. In Proceedings of the 31st ACM International Conference on Multimedia . 17–26
2023
-
[26]
Lei Jin, Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Guannan Jiang, Annan Shu, and Rongrong Ji. 2023. Refclip: A universal teacher for weakly supervised referring expression comprehension. In Proceedings of the IEEE/CVF CVPR . 2681–2690
2023
-
[27]
Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. 2021. Mdetr-modulated detection for end-to-end multi-modal understanding. In Proceedings of the IEEE/CVF ICCV . 1780–1790
2021
-
[28]
Seil Kang, Jinyeong Kim, Junhyeok Kim, and Seong Jae Hwang. 2025. Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding. arXiv preprint arXiv:2503.06287 (2025)
2025 arXiv
-
[29]
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. 2014. Referitgame: Referring to objects in photographs of natural scenes. InProceedings of the 2014 conference on EMNLP . 787–798
2014
-
[30]
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al
-
[31]
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2021. Deduplicating training data makes language models better. arXiv preprint arXiv:2107.06499 (2021)
2021 arXiv
-
[32]
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International conference on machine learning . PMLR, 12888–12900
2022
-
[33]
In Proceedings of the IEEE/CVF ICCV
Segment anything. In Proceedings of the IEEE/CVF ICCV . 4015–4026
-
[34]
Mingchen Li, Mahdi Soltanolkotabi, and Samet Oymak. 2020. Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks. In International conference on artificial intelligence and statistics . 4313
2020
-
[35]
Ziyi Li, Qinye Zhou, Xiaoyun Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie
-
[36]
Junnan Li, Richard Socher, and Steven CH Hoi. 2020. Dividemix: Learning with noisy labels as semi-supervised learning. arXiv preprint arXiv:2002.07394 (2020)
2020 arXiv
-
[37]
Pengyue Lin, Ruifan Li, Yuzhe Ji, Zhihan Yu, Fangxiang Feng, Zhanyu Ma, and Xiaojie Wang. 2024. Triple Alignment Strategies for Zero-shot Phrase Grounding under Weak Supervision. InProceedings of the 32nd ACM International Conference on Multimedia. 4312–4321
2024
-
[38]
Daqing Liu, Hanwang Zhang, Feng Wu, and Zheng-Jun Zha. 2019. Learning to assemble neural module tree networks for visual grounding. In Proceedings of the IEEE/CVF ICCV. 4673–4682
2019
-
[39]
InProceedings of the IEEE/CVF ICCV
Open-vocabulary object segmentation with diffusion models. InProceedings of the IEEE/CVF ICCV . 7667
-
[40]
Yaoyuan Liang, Zhao Yang, Yansong Tang, Jiashuo Fan, Ziran Li, Jingang Wang, Philip HS Torr, and Shao-Lun Huang. 2023. Luna: Language as continuing anchors for referring expression comprehension. In Proceedings of the 31st ACM International Conference on Multimedia . 5174–5184
2023
-
[41]
Yongfei Liu, Bo Wan, Lin Ma, and Xuming He. 2021. Relation-aware instance refinement for weakly supervised visual grounding. InProceedings of the IEEE/CVF CVPR. 5612–5621
2021
-
[42]
Yang Liu, Jiahua Zhang, Qingchao Chen, and Yuxin Peng. 2023. Confidence-aware pseudo-label learning for weakly supervised visual grounding. In Proceedings of the IEEE/CVF ICCV. 2828–2838
2023
-
[43]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruc- tion tuning. Advances in neural information processing systems 36 (2024)
2024
-
[44]
Xihui Liu, Zihao Wang, Jing Shao, Xiaogang Wang, and Hongsheng Li. 2019. Improving referring expression grounding with cross-modal attention-guided erasing. In Proceedings of the IEEE/CVF CVPR . 1950–1959
2019
-
[45]
when to update
Eran Malach and Shai Shalev-Shwartz. 2017. Decoupling “when to update” from “how to update”. Advances in neural information processing systems 30 (2017)
2017
-
[46]
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. 2016. Generation and comprehension of unambiguous object descriptions. In Proceedings of the IEEE CVPR . 11–20. Conference, Zilin Du, Haoxin Li, Jianfei Yu, and Boyang Li
2016
-
[47]
Tao Ma, Bing Bai, Haozhe Lin, Heyuan Wang, Yu Wang, Lin Luo, and Lu Fang
-
[48]
Mansheej Paul, Surya Ganguli, and Gintare Karolina Dziugaite. 2021. Deep learning on a data diet: Finding important examples early in training. Advances in neural information processing systems 34 (2021), 20596–20607
2021
-
[49]
Anas Mahmoud, Mostafa Elhoushi, Amro Abbas, Yu Yang, Newsha Ardalani, Hugh Leather, and Ari S Morcos. 2024. Sieve: Multimodal dataset pruning using image captioning models. In Proceedings of the IEEE/CVF CVPR . 22423–22432
2024
-
[50]
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2023. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952 (2023)
2023 arXiv
-
[51]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF CVPR . 10684–10695
2022
-
[52]
Varun K Nagaraja, Vlad I Morariu, and Larry S Davis. 2016. Modeling context between objects for referring expression understanding. In Computer Vision– ECCV 2016: 14th European Conference . Springer, 792–807
2016
-
[53]
Noam Rotstein, David Bensaïd, Shaked Brody, Roy Ganz, and Ron Kimmel. 2024. Fusecap: Leveraging large language models for enriched fused image captions. In Proceedings of the IEEE/CVF W ACV. 5689–5700
2024
-
[54]
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The RefinedWeb dataset for Falcon LLM: outperforming curated corpora with web data, and web data only.arX...
2023 arXiv
-
[55]
Fengyuan Shi, Ruopeng Gao, Weilin Huang, and Limin Wang. 2023. Dynamic mdetr: A dynamic multimodal transformer decoder for visual grounding. IEEE Transactions on Pattern Analysis and Machine Intelligence (2023)
2023
-
[56]
Hwanjun Song, Minseok Kim, Dongmin Park, and Jae-Gil Lee. 2019. How does early stopping help generalization against label noise? arXiv preprint arXiv:1911.08059 (2019)
2019 arXiv
-
[57]
Daniel Rose, Vaishnavi Himakunthala, Andy Ouyang, Ryan He, Alex Mei, Yujie Lu, Michael Saxon, Chinmay Sonar, Diba Mirza, and William Yang Wang. 2023. Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings.arXiv preprint arXiv:2305.02317 (2023)
2023 arXiv
-
[58]
Wei Su, Peihan Miao, Huanzhang Dou, and Xi Li. 2024. Scanformer: Referring expression comprehension by iteratively scanning. InProceedings of the IEEE/CVF CVPR. 13449–13458
2024
-
[59]
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedan- tam, Devi Parikh, and Dhruv Batra. 2017. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE ICCV . 618–626
2017
-
[60]
Jiamu Sun, Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Guannan Jiang, Zhiyu Wang, and Rongrong Ji. 2023. Refteacher: A strong baseline for semi-supervised referring expression comprehension. In Proceedings of the IEEE/CVF CVPR . 19144–19154
2023
-
[61]
Mingjie Sun, Jimin Xiao, Eng Gee Lim, Si Liu, and John Y Goulermas. 2021. Discriminative triad matching and reconstruction for weakly referring expression grounding. IEEE transactions on pattern analysis and machine intelligence 43, 11 (2021), 4189–4195
2021
-
[62]
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari Morcos
-
[63]
Saksham Suri, Fanyi Xiao, Animesh Sinha, Sean Culatana, Raghuraman Krish- namoorthi, Chenchen Zhu, and Abhinav Shrivastava. 2024. Gen2Det: Generate to Detect. In Synthetic Data for Computer Vision Workshop@ CVPR 2024
2024
-
[64]
Tianyi Tang, Yushuo Chen, Yifan Du, Junyi Li, Wayne Xin Zhao, and Ji-Rong Wen
-
[65]
Wei Su, Peihan Miao, Huanzhang Dou, Gaoang Wang, Liang Qiao, Zheyang Li, and Xi Li. 2023. Language adaptive weight generation for multi-task visual grounding. In Proceedings of the IEEE/CVF CVPR . 10857–10866
2023
-
[66]
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J Gordon. 2018. An empirical study of example forgetting during deep neural network learning. arXiv preprint arXiv:1812.05159 (2018)
2018 arXiv
-
[67]
Shijie Wang, Dahun Kim, Ali Taalimi, Chen Sun, and Weicheng Kuo. 2024. Learn- ing Visual Grounding from Generative Vision and Language Model.arXiv preprint arXiv:2407.14563 (2024)
2024 arXiv
-
[68]
Yuxi Sun, Shanshan Feng, Xutao Li, Yunming Ye, Jian Kang, and Xu Huang
-
[69]
In Proceedings of the 30th ACM International conference on Multimedia
Visual grounding in remote sensing images. In Proceedings of the 30th ACM International conference on Multimedia . 404–412
-
[70]
Dulanga Weerakoon, Vigneshwaran Subbaraju, Tuan Tran, and Archan Misra
-
[71]
Xiaobo Xia, Jiale Liu, Jun Yu, Xu Shen, Bo Han, and Tongliang Liu. 2022. Moderate coreset: A universal method of data selection for real-world data-efficient deep learning. In The Eleventh International Conference on Learning Representations
2022
-
[72]
arXiv preprint arXiv:2305.16944 (2023)
Learning to Imagine: Visually-Augmented Natural Language Generation. arXiv preprint arXiv:2305.16944 (2023)
2023 arXiv
-
[73]
Yonglong Tian, Lijie Fan, Kaifeng Chen, Dina Katabi, Dilip Krishnan, and Phillip Isola. 2023. Learning vision from models rivals learning vision from data. arXiv preprint arXiv:2312.17742 (2023)
2023 arXiv
-
[74]
Linhui Xiao, Xiaoshan Yang, Fang Peng, Yaowei Wang, and Changsheng Xu. 2024. HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual Grounding. arXiv preprint arXiv:2404.13400 (2024)
2024 arXiv
-
[75]
Linhui Xiao, Xiaoshan Yang, Fang Peng, Ming Yan, Yaowei Wang, and Chang- sheng Xu. 2023. CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding. IEEE Transactions on Multimedia (2023)
2023
-
[76]
Sai Wang, Yutian Lin, and Yu Wu. 2024. Omni-Q: Omni-Directional Scene Understanding for Unsupervised Visual Grounding. InProceedings of the IEEE/CVF CVPR. 14261–14270
2024
-
[77]
Yibo Wang, Ruiyuan Gao, Kai Chen, Kaiqiang Zhou, Yingjie Cai, Lanqing Hong, Zhenguo Li, Lihui Jiang, Dit-Yan Yeung, Qiang Xu, et al . 2024. Detdiffusion: Synergizing generative and perceptive models for enhanced data generation and perception. In Proceedings of the IEEE/CVF CV...
2024
-
[78]
Yunyi Xuan, Weijie Chen, Shicai Yang, Di Xie, Luojun Lin, and Yueting Zhuang
-
[79]
In Proceedings of the 30th ACM International Conference on Multimedia
SoftSkip: Empowering multi-modal dynamic pruning for single-stage referring comprehension. In Proceedings of the 30th ACM International Conference on Multimedia. 3608–3616
-
[80]
Siming Yan, Min Bai, Weifeng Chen, Xiong Zhou, Qixing Huang, and Li Erran Li
-
[81]
Kai Xiao, Logan Engstrom, Andrew Ilyas, and Aleksander Madry. 2020. Noise or signal: The role of image backgrounds in object recognition. arXiv preprint arXiv:2006.09994 (2020)
2020 arXiv
-
[82]
Linhui Xiao, Xiaoshan Yang, Xiangyuan Lan, Yaowei Wang, and Changsheng Xu. 2024. Towards Visual Grounding: A Survey. arXiv preprint arXiv:2412.20206 (2024)
2024 arXiv
-
[83]
Shuo Yang, Zeke Xie, Hanyu Peng, Min Xu, Mingming Sun, and Ping Li. 2022. Dataset pruning: Reducing training data by examining generalization influence. arXiv preprint arXiv:2205.09329 (2022)
2022 arXiv
-
[84]
Yue Yang, Wenlin Yao, Hongming Zhang, Xiaoyang Wang, Dong Yu, and Jianshu Chen. 2022. Z-LaVI: Zero-Shot Language Solver Fueled by Visual Imagination. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 1186–1203
2022
-
[85]
Jiahao Xie, Wei Li, Xiangtai Li, Ziwei Liu, Yew Soon Ong, and Chen Change Loy
-
[86]
International Journal of Computer Vision (2024), 1–20
Mosaicfusion: Diffusion models as data augmenters for large vocabulary instance segmentation. International Journal of Computer Vision (2024), 1–20
2024
-
[87]
Sang Michael Xie, Shibani Santurkar, Tengyu Ma, and Percy S Liang. 2023. Data selection for language models via importance resampling. Advances in Neural Information Processing Systems 36 (2023), 34201–34227
2023
-
[88]
Ruilin Yao, Shengwu Xiong, Yichen Zhao, and Yi Rong. 2024. Visual Ground- ing with Multi-modal Conditional Adaptation. In Proceedings of the 32nd ACM International Conference on Multimedia . 3877–3886
2024
-
[89]
In Proceedings of the 31st ACM International Conference on Multimedia
Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt Diversification. In Proceedings of the 31st ACM International Conference on Multimedia. 4928–4938
-
[90]
Han Xue, Zhiwu Huang, Qianru Sun, Li Song, and Wenjun Zhang. 2023. Freestyle layout-to-image synthesis. In Proceedings of the IEEE/CVF CVPR . 14256–14266
2023
-
[91]
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg. 2018. Mattnet: Modular attention network for referring ex- pression comprehension. In Proceedings of the IEEE CVPR . 1307–1315
2018
-
[92]
Vigor: Improving visual grounding of large vision language models with fine-grained reward modeling. In ECCV. Springer, 37–53
-
[93]
Lihe Yang, Xiaogang Xu, Bingyi Kang, Yinghuan Shi, and Hengshuang Zhao. 2024. Freemask: Synthetic images with dense annotations make stronger segmentation models. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[94]
Li Yang, Yan Xu, Chunfeng Yuan, Wei Liu, Bing Li, and Weiming Hu. 2022. Improv- ing visual grounding with visual-linguistic verification and iterative reasoning. In Proceedings of the IEEE/CVF CVPR . 9499–9508
2022
-
[95]
Yunan Zeng, Yan Huang, Jinjin Zhang, Zequn Jie, Zhenhua Chai, and Liang Wang. 2024. Investigating compositional challenges in vision-language models for visual grounding. In Proceedings of the IEEE/CVF CVPR . 14141–14151
2024
-
[96]
Manlin Zhang, Jie Wu, Yuxi Ren, Ming Li, Jie Qin, Xuefeng Xiao, Wei Liu, Rui Wang, Min Zheng, and Andy J Ma. 2023. Diffusionengine: Diffusion model is scalable data engine for object detection. arXiv preprint arXiv:2309.03893 (2023)
2023 arXiv
-
[97]
Ziyan Yang, Kushal Kafle, Franck Dernoncourt, and Vicente Ordonez. 2023. Im- proving visual grounding by encouraging consistent gradient-based explanations. In Proceedings of the IEEE/CVF CVPR . 19165–19174
2023
-
[98]
Zuhao Yang, Fangneng Zhan, Kunhao Liu, Muyu Xu, and Shijian Lu. 2023. AI- Generated Images as Data Source: The Dawn of Synthetic Era. arXiv preprint arXiv:2310.01830 (2023)
2023 arXiv
-
[99]
Quanming Yao, Hansi Yang, Bo Han, Gang Niu, and James Tin-Yau Kwok. 2020. Searching to exploit memorization effect in learning with noisy labels. In Inter- national Conference on Machine Learning . PMLR, 10789–10798
2020
-
[101]
Fulong Ye, Yuxing Long, Fangxiang Feng, and Xiaojie Wang. 2023. Whether you can locate or not? Interactive Referring Expression Generation. In Proceedings of the 31st ACM International Conference on Multimedia . 4697–4706
2023
-
[102]
Jiabo Ye, Junfeng Tian, Ming Yan, Xiaoshan Yang, Xuwu Wang, Ji Zhang, Liang He, and Xin Lin. 2022. Shifting more attention to visual backbone: Query- modulated refinement networks for end-to-end visual grounding. In Proceedings of the IEEE/CVF CVPR . 15502
2022
-
[104]
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg
-
[106]
Xingrui Yu, Bo Han, Jiangchao Yao, Gang Niu, Ivor Tsang, and Masashi Sugiyama
-
[108]
Zhihan Yu and Ruifan Li. 2024. Revisiting counterfactual problems in referring expression comprehension. In Proceedings of the IEEE/CVF CVPR . 13438–13448. Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding Conference,
2024
-
[111]
Haoyu Zhao, Wenhang Ge, and Ying-cong Chen. 2024. LLM-Optic: Unveiling the Capabilities of Large Language Models for Universal Visual Grounding. arXiv preprint arXiv:2405.17104 (2024)
2024 arXiv
-
[112]
Hanqing Zhao, Dianmo Sheng, Jianmin Bao, Dongdong Chen, Dong Chen, Fang Wen, Lu Yuan, Ce Liu, Wenbo Zhou, Qi Chu, et al . 2023. X-paste: Revisiting scalable copy-paste for instance segmentation using clip and stablediffusion. In International Conference on Machine Learning . P...
2023
-
[113]
Minghang Zheng, Jiahua Zhang, Qingchao Chen, Yuxin Peng, and Yang Liu. 2024. ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual Grounding. In Proceedings of the 32nd ACM International Conference on Multimedia. 1187–1196
2024
-
[2016]
In ECCV 2016, Proceedings, Part II 14
Modeling context in referring expressions. In ECCV 2016, Proceedings, Part II 14. Springer, 69–85
2016
-
[2019]
In International Conference on Machine Learning
How does disagreement help generalization against label corruption?. In International Conference on Machine Learning . PMLR, 7164
-
[2021]
In Proceedings of the IEEE/CVF ICCV
Transvg: End-to-end visual grounding with transformers. In Proceedings of the IEEE/CVF ICCV. 1769–1779
-
[2022]
Advances in Neural Information Processing Systems 35 (2022), 19523–19536
Beyond neural scaling laws: beating power law scaling via data pruning. Advances in Neural Information Processing Systems 35 (2022), 19523–19536
2022
-
[2023]
arXiv preprint arXiv:2306.15195 (2023)
Shikra: Unleashing multimodal llm’s referential dialogue magic. arXiv preprint arXiv:2306.15195 (2023)
2023 arXiv
-
[2024]
In Proceedings of the IEEE/CVF CVPR
When visual grounding meets gigapixel-level large-scale scenes: benchmark and approach. In Proceedings of the IEEE/CVF CVPR . 22119–22128
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.