REVIEW 3 major objections 7 minor 113 references
Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation
T0 review · 3 major / 7 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read Mosaic3D claims the largest open-vocabulary 3D mask-text dataset—5.6M pairs across 29K scenes—and uses it to reach state-of-the-art open-vocabulary 3D semantic and instance segmentation.
desk verdict A genuinely large 3D mask-text dataset and a solid model, but the OV3D baseline and the missing human evaluation of data quality need fixing before the headline claims are fully trustworthy. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the data engine: RAM++ tags objects, Grounding-DINO proposes boxes, SAM2 and SEEM produce precise foreground and panoptic masks, Osprey writes region-specific captions, and a projection-plus-depth-inclusion test transfers each 2D mask onto 3D points. On the model side, the contrastive per-point loss aligns point features with text embeddings, and the mask decoder's caption loss aligns mask embeddings with captions, supporting open-vocabulary instance segmentation without ground-truth labels.
What would settle it
Take a random sample of Mosaic3D-5.6M mask-text pairs, show a human the masked RGB region and the caption, and measure agreement; if captions frequently describe content outside the mask, the dataset's core premise fails. A cheaper quantitative probe: retrain the encoder with Osprey captions replaced by class-name-only labels from RAM++, and check whether the ScanNet200 gains collapse; if they do not collapse, the detailed region captions are not carrying the claimed benefit.
Extended reading notes
Core claim
The paper claims that large-scale, precisely masked, richly captioned 3D training data can be produced automatically—no human annotation—by combining open-vocabulary image segmentation (Grounded-SAM and SEEM) with a region-aware vision-language model (Osprey) and projecting the resulting 2D masks to 3D points through a depth inclusion test. On this data, contrastive per-point training aligns 3D geometry with text embeddings, and a Mask3D-style decoder, trained with an added caption loss, turns those aligned features into open-vocabulary instance predictions. Training on the full 5.6M-pair dataset yields state-of-the-art f-mIoU across four semantic segmentation benchmarks and a single-stage 3D-only instance segmenter that runs in about one second per scene, in contrast to prior methods that require multi-view CLIP inference.
Load-bearing premise
The whole dataset's value rests on the 2D teachers: if Osprey's captions do not actually describe the masked region, or if the projection and inclusion test attaches masks to the wrong 3D points, then the scale and apparent gains are artifacts of the teachers rather than genuine 3D understanding.
Editorial extensions
If this is right
- If the data engine transfers to new scene datasets, scaling 3D open-vocabulary understanding becomes a matter of running 2D foundation models on more RGB-D scans rather than hiring annotators.
- The single-stage 3D instance segmentation model shows that language-aligned 3D features can carry open-vocabulary instance prediction directly, making multi-view CLIP inference at test time unnecessary.
- The monotonic gains as datasets are added suggest that further scaling of mask-text pairs will keep improving open-vocabulary semantic segmentation.
- The zero-shot results with anonymized class names suggest that the model learns region semantics beyond memorized class labels, which is closer to true open-vocabulary behavior.
- The caption-loss-aligned mask decoder could be reused as a proposal-free 3D vision-language interface for referring segmentation and other language-grounded tasks.
Reading between the lines
- If the pipeline's quality holds, the same teacher-student recipe could extend to outdoor, dynamic, or object-centric 3D data without new annotation, since the projection step only needs posed RGB-D frames.
- The paper's own zero-shot anonymization experiment implies that part of previous methods' apparent success came from class names leaking into training captions; Mosaic3D-5.6M appears less reliant on that leak, but the residual drop still leaves much of the open-vocabulary gain tied to caption content.
- A direct human audit of a random sample of mask-caption pairs would be a cheap, decisive test of whether Osprey's descriptions are grounded in the masked region rather than hallucinated from context.
- Combining the mask decoder with a proposal method like Segment3D suggests a modular path: any improved 3D proposal network could slot in and lift instance segmentation further.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces an automatic data-generation pipeline that combines Grounded-SAM2/SEEM mask proposals with the Osprey region-aware VLM to create 3D mask-text pairs, and applies it to ScanNet, ARKitScenes, Matterport3D, ScanNet++, and Structured3D to construct Mosaic3D-5.6M (about 5.6M captions across about 29K scenes). It then trains a SparseUNet encoder with a point-text contrastive loss and a Mask3D-style decoder with a mask-caption loss, reporting state-of-the-art open-vocabulary semantic segmentation on ScanNet20, ScanNet200, Matterport3D, and ScanNet++, plus a single-stage open-vocabulary 3D instance segmentation method that does not require ground-truth labels. The paper includes data-scaling and model-scaling studies, component ablations, a class-name anonymization experiment, and an extensive supplementary appendix.
Significance. If the claims hold, Mosaic3D-5.6M would be a valuable community asset: it is substantially larger than existing mask-text datasets, the pipeline is automatic, and the authors provide a project page and extensive ablations that isolate the contributions of mask generators, captioners, frame sampling, data scale, model scale, and text encoders. The anonymization study is a useful attempt to separate class-name memorization from open-vocabulary generalization. The main weaknesses are the absence of direct human validation of the generated captions and masks, the internally inconsistent OV3D baseline, and the lack of clarity on whether evaluation scenes were excluded from the generated training data; these issues must be resolved before the central benchmark claims can be accepted.
major comments (3)
- [Section 3.2 / Table A1 / Section 5.3 / Table 1] The paper never states that the data-generation pipeline was restricted to the official training splits of the source datasets. Table A1 reports 1,513 ScanNet scenes, 2,194 Matterport3D scenes, and 380 ScanNet++ scenes, which appear to be the full datasets, while Table 1 evaluates Mosaic3D-5.6M on ScanNet20/200 validation, Matterport3D test, and ScanNet++ validation. If captions from evaluation scenes were included in training, the benchmark results are contaminated by scene-level leakage, because the model has seen the same point clouds and their captions during training. Please specify the exact train/validation/test split used for each source dataset, state whether any evaluation scene appears in Mosaic3D-5.6M, and re-run the affected benchmarks if leakage occurred.
- [Section 3.2 / Table A2 / Figure 2] The 'highest-quality dataset' claim is not directly validated. The metrics reported in Table A2, namely noun count, coverage, and entropy, cannot detect hallucinated captions or cross-view mask misalignment, and mask entropy rewards a partial mask covering a subset of one ground-truth instance as much as a complete mask of that instance, since both have low entropy. No human evaluation of caption correctness or mask-boundary quality is reported in the main text or the supplementary. Please add a human study on a random sample of mask-text pairs, reporting e.g. caption relevance/accuracy and mask IoU against ground-truth or manually refined masks, and state how the qualitative examples in Figure A2 were selected.
- [Table 1 / Appendix B.3 / Table A7] The OV3D comparison is internally inconsistent. Table 1 reports the original OV3D numbers of 64.0 f-mIoU on ScanNet20 and 8.7 on ScanNet200, while Table A7 shows that the authors' own reimplementation, OV3D-rep with DenseAlign, obtains only 34.7/4.6, and the improved OV3D++ with Contrastive obtains 58.4/9.2. The text's claim of 'surpassing OV3D by 1.0p' on ScanNet20 therefore depends on a baseline that the paper itself cannot reproduce. Please present the original-paper number and the same-protocol reproduction, ideally OV3D++ Contrastive under the shared SPUNet34C/Recap-CLIP setting, side by side in Table 1, and make the claims in the text consistent with the chosen baseline.
minor comments (7)
- [Abstract / Section 3.2 / Table A1] The abstract and Section 3.2 say 'over 30K scenes', but Table A1 sums to 29,197 scenes; please correct the count or rephrase to 'about 29K scenes'.
- [Section 3.2 / Table A1] Section 3.2 states 'approximately 1M RGB-D frames', while Table A1 reports a total of 7.1M frames; please reconcile these numbers.
- [Figure 2 / Table A2] Figure 2 reports Entropy 60.7 for Mosaic3D-5.6M, but Table A2 reports the same value for Mosaic3D-SN and leaves the entropy blank for Mosaic3D-5.6M; the figure caption should state the subset used for the entropy statistic.
- [Section 3.1 / Appendix A.3 / Table 3] The main text says 'Grounded-SAM' for mask generation, while Appendix A.3 and Table 3 use Grounded-SAM2 with SAM2 checkpoints; please unify the terminology throughout the paper.
- [Section 4.2 / Equation (5)] Equation (5) uses the subscript k in the numerator text embedding \bar ztext_k after defining the mask embeddings with subscript m; replace k with m for consistency.
- [Section 3.1 / Equation (1)] Equation (1) does not state that the projected pixel must lie inside the image bounds, and no sensitivity analysis is given for the depth threshold epsilon; please clarify the bounds check and, ideally, include an ablation over epsilon.
- [Tables 1-4] All benchmark tables report a single training run; given margins as small as 1.0 f-mIoU in Table 1, please report at least three seeds with mean and standard deviation for the key comparisons.
Circularity Check
No significant circularity: benchmarks are external, and the anonymization experiment directly breaks the class-name loop.
full rationale
The paper's derivation chain is: (i) generate 3D mask-text pairs from 2D open-vocabulary segmentation models and a region-aware VLM via projection and inclusion test (Eq. 1); (ii) train a 3D encoder with a contrastive loss against a frozen text encoder (Eq. 2); (iii) train a mask decoder using caption-merged Segment3D masks; and (iv) evaluate on external benchmarks (ScanNet, ScanNet++, Matterport3D, ScanNet200). No equation reduces to a fitted parameter or to the benchmark numbers. The training objective does align point features to CLIP text embeddings, and evaluation also uses CLIP text embeddings of class names, but that is the definition of open-vocabulary segmentation, not a forced reduction: the 3D encoder could fail to learn the alignment, and the benchmarks provide independent human-annotated labels. The paper explicitly addresses the concern that captions contain evaluation class names with an anonymization experiment (Table 4), replacing class names with 'object' and showing that Mosaic3D-5.6M still retains the strongest performance; this is a direct control against the main potential circularity. The dataset-quality claim relies on the accuracy of teacher models (Grounded-SAM, SEEM, Osprey) without human evaluation of the generated pairs, and the mask-entropy metric in Table A2 measures homogeneity against GT instance IDs rather than caption correctness. These are validation gaps and correctness risks, not circularity by construction. Self-citations (e.g., MinkowskiNet for the sparse-conv backbone) are standard architectural references and are not load-bearing for the paper's central data-scaling or state-of-the-art claims. Consequently, the analysis finds no significant circularity.
Assumptions & free parameters
free parameters (5)
- Depth threshold epsilon in Eq. (1) =
Not reported
- IoU threshold tau in Algorithm 1 =
Not reported
- Number of frames per scene K =
25 or 125 in ablations
- Loss weights lambda_obj, lambda_dice, lambda_bce, lambda_cap =
2, 5, 2, 1
- Grounding-DINO thresholds and NMS =
box score 0.25, text score 0.2, NMS IoU 0.5, max box area 95%
assumptions (5)
- domain assumption 2D segmentation models Grounded-SAM and SEEM produce accurate open-vocabulary masks for the pipeline.
- domain assumption Osprey region captioning produces accurate, diverse, and contextually correct captions for masked regions.
- domain assumption The projection and inclusion test in Eq. (1) correctly associates 2D masks to 3D points given accurate camera poses and depth.
- domain assumption CLIP or Recap-CLIP text embeddings adequately represent caption semantics for contrastive learning.
- domain assumption Segment3D class-agnostic masks provide a good proposal set for instance segmentation training.
Cite this review
Pith. "Pith review of Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation." pith.science (2026). https://pith.science/paper/QVN2ADES
@misc{pith2026250202548,
author = {Pith},
title = {Pith review of: Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/QVN2ADES}},
note = {Machine review of arXiv:2502.02548}
}
read the original abstract
We tackle open-vocabulary 3D scene understanding by introducing a novel data generation pipeline and training framework. Our method addresses three critical requirements for effective training: precise 3D region segmentation, comprehensive textual descriptions, and sufficient dataset scale. By leveraging state-of-the-art open-vocabulary image segmentation models and region-aware Vision-Language Models, we develop an automatic pipeline that generates high-quality 3D mask-text pairs. Applying this pipeline to multiple 3D scene datasets, we create Mosaic3D-5.6M, a dataset of over 30K annotated scenes with 5.6M mask-text pairs, significantly larger than existing datasets. Building upon this data, we propose Mosaic3D, a foundation model combining a 3D encoder trained with contrastive learning and a lightweight mask decoder for open-vocabulary 3D semantic and instance segmentation. Our approach achieves state-of-the-art results on open-vocabulary 3D semantic and instance segmentation tasks including ScanNet200, Matterport3D, and ScanNet++, with ablation studies validating the effectiveness of our large-scale training data.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
https://huggingface
Vit-gpt2 image captioning. https://huggingface. co / nlpconnect / vit - gpt2 - image - captioning, 2022. 2, 3
2022
-
[2]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ah- mad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774,
-
[3]
Panos Achlioptas, Ahmed Abdelreheem, Fei Xia, Mohamed Elhoseiny, and Leonidas J. Guibas. ReferIt3D: Neural lis- teners for fine-grained 3d object identification in real-world scenes. In 16th European Conference on Computer Vision (ECCV), 2020. 2
2020
-
[4]
Scanqa: 3d question answering for spatial scene understanding
Daichi Azuma, Taiki Miyanishi, Shuhei Kurita, and Motoaki Kawanabe. Scanqa: 3d question answering for spatial scene understanding. In proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 19129– 19139, 2022. 2
2022
-
[5]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023. 2
arXiv 2023
-
[6]
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond. arXiv preprint arXiv:2308.12966, 2023. 2
arXiv 2023
-
[7]
Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data
Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, Yuri Fei- gin, Peter Fu, Thomas Gebauer, Daniel Kurz, Tal Dimry, Brandon Joffe, Arik Schwartz, et al. Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1), ...
2021
-
[8]
Audiolm: A language modeling approach to audio generation
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Rob- lek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al. Audiolm: A language modeling approach to audio generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:2523–2533, 2023. 2
2023
Show all 113 references
-
[9]
Large-scale machine learning with stochas- tic gradient descent
Léon Bottou. Large-scale machine learning with stochas- tic gradient descent. In Proceedings of COMPSTAT’2010: 19th International Conference on Computational Statistic- sParis France, August 22-27, 2010 Keynote, Invited and Contributed Papers, pages 177–186. Springer, 2010. 6
2010
-
[10]
Language models are few-shot learners
Tom B Brown. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020. 2
2005 arXiv
-
[11]
Coyo-700m: Image-text pair dataset
Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim. Coyo-700m: Image-text pair dataset. https : / / github . com / kakaobrain/coyo-dataset, 2022. 1
2022
-
[12]
End- to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End- to-end object detection with transformers. In European conference on computer vision , pages 213–229. Springer,
-
[13]
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jé- gou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021. 2
2021
-
[14]
Matterport3d: Learning from rgb-d data in indoor environments
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Hal- ber, Matthias Niebner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments. In 2017 International Confer- ence on 3D Vision (3DV), pages 667–676. IEEE Comp...
2017
-
[15]
Scanrefer: 3d object localization in rgb-d scans using natural language
Dave Zhenyu Chen, Angel X Chang, and Matthias Nießner. Scanrefer: 3d object localization in rgb-d scans using natural language. In European conference on computer vision, pages 202–221. Springer, 2020. 2, 4, 5
2020
-
[16]
Pali: A jointly-scaled multilingual language-image model
Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergio- vanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. Pali: A jointly-scaled multilingual language-image model. In The Eleventh International Conference on Learning Represent...
2022
-
[17]
Per-pixel classification is not all you need for semantic segmentation
Bowen Cheng, Alexander G Schwing, and Alexander Kir- illov. Per-pixel classification is not all you need for semantic segmentation. In 35th Conference on Neural Information Processing Systems, NeurIPS 2021 , pages 17864–17875. Neural information processing systems foundation, ...
2021
-
[18]
Masked-attention mask transformer for universal image segmentation
Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexan- der Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1290–1299, 2022. 2, 5
2022
-
[19]
Gonzalez, Ion Stoica, and Eric P
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023. 2
2023
-
[20]
Cat- seg: Cost aggregation for open-vocabulary semantic seg- mentation
Seokju Cho, Heeseong Shin, Sunghwan Hong, Anurag Arnab, Paul Hongsuck Seo, and Seungryong Kim. Cat- seg: Cost aggregation for open-vocabulary semantic seg- mentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4113–4123,
-
[21]
4d spatio-temporal convnets: Minkowski convolutional neural networks
Christopher Choy, JunYoung Gwak, and Silvio Savarese. 4d spatio-temporal convnets: Minkowski convolutional neural networks. In Proceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition , pages 3075–3084,
-
[22]
Scaling instruction- finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instruction- finetuned language models. Journal of Machine Learning Research, 25(70):1–53, 2024. 2
2024
-
[23]
Pointcept: A codebase for point cloud perception research
Pointcept Contributors. Pointcept: A codebase for point cloud perception research. https://github.com/ Pointcept/Pointcept, 2023. 2
2023
-
[24]
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5828–5839, 2017. 1, 2, 4, ...
2017
-
[25]
Procthor: Large-scale embodied ai using procedural generation
Matt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs, Kiana Ehsani, Jordi Salvador, Winson Han, Eric Kolve, Aniruddha Kembhavi, and Roozbeh Mottaghi. Procthor: Large-scale embodied ai using procedural generation. Ad- vances in Neural Information Processing Systems, 35:5982...
2022
-
[26]
Pengi: An audio language model for audio tasks
Soham Deshmukh, Benjamin Elizalde, Rita Singh, and Huaming Wang. Pengi: An audio language model for audio tasks. Advances in Neural Information Processing Systems, 36:18090–18108, 2023. 2
2023
-
[27]
Pla: Language-driven open- vocabulary 3d scene understanding
Runyu Ding, Jihan Yang, Chuhui Xue, Wenqing Zhang, Song Bai, and Xiaojuan Qi. Pla: Language-driven open- vocabulary 3d scene understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023. 2, 3, 5, 6, 7
2023
-
[28]
The llama 3 herd of models
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783,
-
[29]
Efficient graph-based image segmentation
Pedro F Felzenszwalb and Daniel P Huttenlocher. Efficient graph-based image segmentation. International journal of computer vision, 59:167–181, 2004. 7
2004
-
[30]
Dat- acomp: In search of the next generation of multimodal datasets
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, et al. Dat- acomp: In search of the next generation of multimodal datasets. Advances in Neural Information Processing Sys...
2024
-
[31]
Scal- ing open-vocabulary image segmentation with image-level labels
Golnaz Ghiasi, Xiuye Gu, Yin Cui, and Tsung-Yi Lin. Scal- ing open-vocabulary image segmentation with image-level labels. In European Conference on Computer Vision, pages 540–557. Springer, 2022. 1, 2
2022
-
[32]
Imagebind one embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. Imagebind one embedding space to bind them all. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, ...
2023
-
[33]
3d semantic segmentation with submanifold sparse convolutional networks
Benjamin Graham, Martin Engelcke, and Laurens Van Der Maaten. 3d semantic segmentation with submanifold sparse convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 9224–9232, 2018. 5, 6
2018
-
[34]
Open- vocabulary object detection via vision and language knowl- edge distillation
Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui. Open- vocabulary object detection via vision and language knowl- edge distillation. In International Conference on Learning Representations, 2021. 1
2021
-
[35]
Regiongpt: Towards region understanding vision lan- guage model
Qiushan Guo, Shalini De Mello, Hongxu Yin, Wonmin Byeon, Ka Chun Cheung, Yizhou Yu, Ping Luo, and Sifei Liu. Regiongpt: Towards region understanding vision lan- guage model. arXiv preprint arXiv:2403.02330, 2024. 2
2024 arXiv
-
[36]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 3
2016
-
[37]
Denoising diffu- sion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. 2
2020
-
[38]
Scaling up vision-language pre-training for image captioning
Xiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang, Zicheng Liu, Yumao Lu, and Lijuan Wang. Scaling up vision-language pre-training for image captioning. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 17980–17989, 2022. 1
2022
-
[39]
An Embodied Generalist Agent in 3D World, 2023
Jiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu, Puhao Li, Yan Wang, Qing Li, Song-Chun Zhu, Baoxiong Jia, and Siyuan Huang. An Embodied Generalist Agent in 3D World, 2023. 8
2023
-
[40]
Segment3d: Learning fine-grained class-agnostic 3d segmentation without manual labels
Rui Huang, Songyou Peng, Ayca Takmaz, Federico Tombari, Marc Pollefeys, Shiji Song, Gao Huang, and Francis Engel- mann. Segment3d: Learning fine-grained class-agnostic 3d segmentation without manual labels. In European Confer- ence on Computer Vision, 2024. 2, 5, 7
2024
-
[41]
Open-set image tagging with multi-grained text supervision
Xinyu Huang, Yi-Jie Huang, Youcai Zhang, Weiwei Tian, Rui Feng, Yuejie Zhang, Yanchun Xie, Yaqian Li, and Lei Zhang. Open-set image tagging with multi-grained text supervision. arXiv e-prints, pages arXiv–2310, 2023. 2, 3, 8, 4
2023
-
[42]
Openins3d: Snap and lookup for 3d open-vocabulary instance segmentation
Zhening Huang, Xiaoyang Wu, Xi Chen, Hengshuang Zhao, Lei Zhu, and Joan Lasenby. Openins3d: Snap and lookup for 3d open-vocabulary instance segmentation. In European Conference on Computer Vision, 2024. 2, 5, 6, 7
2024
-
[43]
Oneformer: One transformer to rule universal image segmentation
Jitesh Jain, Jiachen Li, Mang Tik Chiu, Ali Hassani, Nikita Orlov, and Humphrey Shi. Oneformer: One transformer to rule universal image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2989–2998, 2023. 2
2023
-
[44]
Scen- eVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding, 2024
Baoxiong Jia, Yixin Chen, Huangyue Yu, Yan Wang, Xuesong Niu, Tengyu Liu, Qing Li, and Siyuan Huang. Scen- eVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding, 2024. 2, 6, 7, 8
2024
-
[45]
Scaling up visual and vision-language representa- tion learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representa- tion learning with noisy text supervision. In International conference on machine learning, pages 4904–4916. PMLR,
-
[46]
Mistral 7b
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lam- ple, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023. 2
-
[47]
Mixtral of experts
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Deven- dra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024. 2
2024 arXiv
-
[48]
Open-vocabulary 3d semantic segmentation with foundation models
Li Jiang, Shaoshuai Shi, and Bernt Schiele. Open-vocabulary 3d semantic segmentation with foundation models. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21284–21294, 2024. 2, 3, 6, 7, 8, 1, 4
2024
-
[49]
In defense of lazy visual grounding for open-vocabulary semantic segmentation
Dahyun Kang and Minsu Cho. In defense of lazy visual grounding for open-vocabulary semantic segmentation. In European Conference on Computer Vision, 2024. 2
2024
-
[50]
Segment any- thing
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment any- thing. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 4015–4026, 202...
2023
-
[51]
Language-driven semantic seg- mentation
Boyi Li, Kilian Q Weinberger, Serge Belongie, Vladlen Koltun, and Rene Ranftl. Language-driven semantic seg- mentation. In International Conference on Learning Repre- sentations, 2022. 1, 2
2022
-
[52]
Semantic-sam: Segment and recognize anything at any gran- ularity
Feng Li, Hao Zhang, Peize Sun, Xueyan Zou, Shilong Liu, Jianwei Yang, Chunyuan Li, Lei Zhang, and Jianfeng Gao. Semantic-sam: Segment and recognize anything at any gran- ularity. In European Conference on Computer Vision, 2024. 2
2024
-
[53]
Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi. BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, pages 128...
2022
-
[54]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA , pages 19730...
2023
-
[55]
What if we recaption billions of web images with llama-3? arXiv preprint arXiv:2406.08478,
Xianhang Li, Haoqin Tu, Mude Hui, Zeyu Wang, Bingchen Zhao, Junfei Xiao, Sucheng Ren, Jieru Mei, Qing Liu, Huangjie Zheng, et al. What if we recaption billions of web images with llama-3? arXiv preprint arXiv:2406.08478,
-
[56]
Open-vocabulary semantic segmentation with mask-adapted clip
Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, and Diana Marculescu. Open-vocabulary semantic segmentation with mask-adapted clip. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pag...
2023
-
[57]
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26296–26306, 2024. 8, 2
2024
-
[58]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In Advances in neural information processing systems, 2024. 2, 3, 7
2024
-
[59]
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499, 2023. 2, 3
2023 arXiv
-
[60]
Mmscan: A multi-modal 3d scene dataset with hierarchical grounded language annotations
Ruiyuan Lyu, Jingli Lin, Tai Wang, Xiaohan Mao, Yilun Chen, Runsen Xu, Haifeng Huang, Chenming Zhu, Dahua Lin, and Jiangmiao Pang. Mmscan: A multi-modal 3d scene dataset with hierarchical grounded language annotations. Advances in Neural Information Processing Systems , 37: 50...
2024
-
[61]
Multiscan: Scalable rgbd scanning for 3d environments with articulated objects
Yongsen Mao, Yiming Zhang, Hanxiao Jiang, Angel Chang, and Manolis Savva. Multiscan: Scalable rgbd scanning for 3d environments with articulated objects. Advances in neural information processing systems, 35:9058–9071, 2022. 7
2022
-
[62]
V-net: Fully convolutional neural networks for volumetric medical image segmentation
Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. V-net: Fully convolutional neural networks for volumetric medical image segmentation. In 2016 fourth international conference on 3D vision (3DV), pages 565–571. Ieee, 2016. 6
2016
-
[63]
Silc: Improving vision language pretraining with self-distillation
Muhammad Ferjad Naeem, Yongqin Xian, Xiaohua Zhai, Lukas Hoyer, Luc Van Gool, and Federico Tombari. Silc: Improving vision language pretraining with self-distillation. In European Conference on Computer Vision, 2024. 2
2024
-
[64]
Isbnet: a 3d point cloud instance segmentation network with instance- aware sampling and box-aware dynamic convolution
Tuan Duc Ngo, Binh-Son Hua, and Khoi Nguyen. Isbnet: a 3d point cloud instance segmentation network with instance- aware sampling and box-aware dynamic convolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13550–13559, 2023. 7
2023
-
[65]
Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance
Phuc Nguyen, Tuan Duc Ngo, Evangelos Kalogerakis, Chuang Gan, Anh Tran, Cuong Pham, and Khoi Nguyen. Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 4018–402...
2024
-
[66]
Improved denoising diffusion probabilistic models
Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In International conference on machine learning, pages 8162–8171. PMLR,
-
[67]
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel HAZIZA, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rab...
2024
-
[68]
Openscene: 3d scene understanding with open vocabularies
Songyou Peng, Kyle Genova, Chiyu "Max" Jiang, An- drea Tagliasacchi, Marc Pollefeys, and Thomas Funkhouser. Openscene: 3d scene understanding with open vocabularies. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2, 6, 7, 5
2023
-
[69]
Kosmos-2: Ground- ing multimodal large language models to the world
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. Kosmos-2: Ground- ing multimodal large language models to the world. arXiv preprint arXiv:2306.14824, 2023. 2, 3, 8
2023 arXiv
-
[70]
Language models are un- supervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are un- supervised multitask learners. OpenAI blog, 1(8):9, 2019. 2
2019
-
[71]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of th...
2021
-
[72]
Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai
Santhosh Kumar Ramakrishnan, Aaron Gokaslan, Erik Wi- jmans, Oleksandr Maksymets, Alexander Clegg, John M Turner, Eric Undersander, Wojciech Galuba, Andrew West- bury, Angel X Chang, et al. Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai....
2021
-
[73]
Denseclip: Language-guided dense prediction with context- aware prompting
Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, and Jiwen Lu. Denseclip: Language-guided dense prediction with context- aware prompting. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition, pages 1808...
2022
-
[74]
Glamm: Pixel grounding large multimodal model
Hanoona Rasheed, Muhammad Maaz, Sahal Shaji, Abdel- rahman Shaker, Salman Khan, Hisham Cholakkal, Rao M Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S Khan. Glamm: Pixel grounding large multimodal model. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Patter...
2024
-
[75]
Sam 2: Segment anything in images and videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, et al. Sam 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714, 2024. 2, 3, 4, 8
2024 arXiv
-
[76]
Grounded sam: Assembling open-world models for diverse visual tasks
Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao, Jiayu Chen, Xinyu Huang, Yukang Chen, Feng Yan, et al. Grounded sam: Assembling open-world models for diverse visual tasks. arXiv preprint arXiv:2401.14159,
-
[77]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 2
2022
-
[78]
Language- grounded indoor 3d semantic segmentation in the wild
David Rozenberszki, Or Litany, and Angela Dai. Language- grounded indoor 3d semantic segmentation in the wild. In European Conference on Computer Vision, pages 125–141. Springer, 2022. 6, 7, 8, 2, 4, 5
2022
-
[79]
Audiopalm: A large language model that can speak and listen
Paul K Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, et al. Audiopalm: A large language model that can speak and listen. arXiv preprint arXiv:2306.12925,
-
[80]
A multi-view stereo benchmark with high- resolution images and multi-camera videos
Thomas Schops, Johannes L Schonberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and An- dreas Geiger. A multi-view stereo benchmark with high- resolution images and multi-camera videos. In CVPR, 2017. 8
2017
-
[81]
Laion-5b: An open large-scale dataset for train- ing next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, et al. Laion-5b: An open large-scale dataset for train- ing next generation image-text models. Advances in Neural Inf...
2022
-
[82]
Mask3d: Mask trans- former for 3d semantic instance segmentation
Jonas Schult, Francis Engelmann, Alexander Hermans, Or Litany, Siyu Tang, and Bastian Leibe. Mask3d: Mask trans- former for 3d semantic instance segmentation. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 8216–8223. IEEE, 2023. 2, 5, 7
2023
-
[83]
Clip-fields: Weakly supervised semantic fields for robotic memory
Nur Muhammad Mahi Shafiullah, Chris Paxton, Lerrel Pinto, Soumith Chintala, and Arthur Szlam. Clip-fields: Weakly supervised semantic fields for robotic memory. InICRA2023 Workshop on Pretraining for Robotics (PT4R), 2023. 1
2023
-
[84]
Super-convergence: Very fast training of neural networks using large learning rates
Leslie N Smith and Nicholay Topin. Super-convergence: Very fast training of neural networks using large learning rates. In Artificial intelligence and machine learning for multi-domain operations applications, pages 369–386. SPIE,
-
[85]
Denois- ing diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denois- ing diffusion implicit models. In International Conference on Learning Representations, 2021. 2
2021
-
[86]
Score-based generative modeling through stochastic differential equa- tions
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equa- tions. In International Conference on Learning Representa- tions, 2021. 2
2021
-
[87]
Open- mask3d: open-vocabulary 3d instance segmentation
Ayça Takmaz, Elisabetta Fedele, Robert W Sumner, Marc Pollefeys, Federico Tombari, and Francis Engelmann. Open- mask3d: open-vocabulary 3d instance segmentation. In Proceedings of the 37th International Conference on Neural Information Processing Systems, pages 68367–68390, 20...
2023
-
[88]
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023. 2
2023 arXiv
-
[89]
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of con- text
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of con- text. arXiv preprint arXiv:2403.05530, 2024. 2
2024 arXiv
-
[90]
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Mar- tinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. 2
2023 arXiv
-
[91]
Llama 2: Open foundation and fine-tuned chat models.arXiv preprint arXiv:2307.09288, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models.arXiv preprint arXiv:2307.09288, 2023. 2
2023 arXiv
-
[92]
Rio: 3d object instance re- localization in changing indoor environments
Johanna Wald, Armen Avetisyan, Nassir Navab, Federico Tombari, and Matthias Nießner. Rio: 3d object instance re- localization in changing indoor environments. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision, pages 7658–7667, 2019. 7
2019
-
[93]
Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learn- ing framework
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learn- ing framework. In International Conference on Machine Learni...
2022
-
[94]
EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI, 2023
Tai Wang, Xiaohan Mao, Chenming Zhu, Runsen Xu, Ruiyuan Lyu, Peisen Li, Xiao Chen, Wenwei Zhang, Kai Chen, Tianfan Xue, Xihui Liu, Cewu Lu, Dahua Lin, and Jiangmiao Pang. EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI, 2023. 2, 8
2023
-
[95]
Towards large- scale 3d representation learning with multi-dataset point prompt training
Xiaoyang Wu, Zhuotao Tian, Xin Wen, Bohao Peng, Xihui Liu, Kaicheng Yu, and Hengshuang Zhao. Towards large- scale 3d representation learning with multi-dataset point prompt training. arXiv preprint arXiv:2308.09718, 2023. 6, 4
2023 arXiv
-
[96]
Groupvit: Semantic segmentation emerges from text supervision
Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang. Groupvit: Semantic segmentation emerges from text supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18134–18144, 2022. 1
2022
-
[97]
Qwen2 technical report.arXiv preprint arXiv:2407.10671, 2024
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al. Qwen2 technical report.arXiv preprint arXiv:2407.10671, 2024. 2
2024 arXiv
-
[98]
Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding
Jihan Yang, Runyu Ding, Weipeng Deng, Zhe Wang, and Xi- aojuan Qi. Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024. 2, 3, 5, 6, 7, 8
2024
-
[99]
Scannet++: A high-fidelity dataset of 3d in- door scenes
Chandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, and Angela Dai. Scannet++: A high-fidelity dataset of 3d in- door scenes. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12–22, 2023. 1, 2, 4, 6, 7, 5
2023
-
[100]
Sai3d: Segment any instance in 3d scenes
Yingda Yin, Yuzheng Liu, Yang Xiao, Daniel Cohen-Or, Jingwei Huang, and Baoquan Chen. Sai3d: Segment any instance in 3d scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3292–3302, 2024. 5, 7
2024
-
[101]
Ferret: Refer and ground anything anywhere at any granularity
Haoxuan You, Haotian Zhang, Zhe Gan, Xianzhi Du, Bowen Zhang, Zirui Wang, Liangliang Cao, Shih-Fu Chang, and Yinfei Yang. Ferret: Refer and ground anything anywhere at any granularity. arXiv preprint arXiv:2310.07704, 2023. 2, 3, 8
2023 arXiv
-
[102]
Convolutions die hard: open-vocabulary seg- mentation with single frozen convolutional clip
Qihang Yu, Ju He, Xueqing Deng, Xiaohui Shen, and Liang- Chieh Chen. Convolutions die hard: open-vocabulary seg- mentation with single frozen convolutional clip. In Pro- ceedings of the 37th International Conference on Neural Information Processing Systems, pages 32215–32234, 2023. 2
2023
-
[103]
Osprey: Pixel understanding with visual instruction tuning
Yuqian Yuan, Wentong Li, Jian Liu, Dongqi Tang, Xinjie Luo, Chi Qin, Lei Zhang, and Jianke Zhu. Osprey: Pixel understanding with visual instruction tuning. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 28202–28211, 2024. 2, 3, 4, 8
2024
-
[104]
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11975–11986, 2023. 1, 4
2023
-
[105]
Speechgpt: Empow- ering large language models with intrinsic cross-modal con- versational abilities
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. Speechgpt: Empow- ering large language models with intrinsic cross-modal con- versational abilities. In The 2023 Conference on Empirical Methods in Natural Language Processing, 2023. 2
2023
-
[106]
Recognize anything: A strong image tagging model
Youcai Zhang, Xinyu Huang, Jinyu Ma, Zhaoyang Li, Zhaochuan Luo, Yanchun Xie, Yuzhuo Qin, Tong Luo, Yaqian Li, Shilong Liu, et al. Recognize anything: A strong image tagging model. arXiv preprint arXiv:2306.03514 ,
-
[107]
Structured3d: A large photo-realistic dataset for structured 3d modeling
Jia Zheng, Junfei Zhang, Jing Li, Rui Tang, Shenghua Gao, and Zihan Zhou. Structured3d: A large photo-realistic dataset for structured 3d modeling. In Computer Vision– ECCV 2020: 16th European Conference, Glasgow, UK, Au- gust 23–28, 2020, Proceedings, Part IX 16, pages 519–53...
2020
-
[108]
Detecting twenty-thousand classes using image-level supervision
Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Krähen- bühl, and Ishan Misra. Detecting twenty-thousand classes using image-level supervision. In European Conference on Computer Vision, pages 350–368. Springer, 2022. 8
2022
-
[109]
Generalized decoding for pixel, image, and language
Xueyan Zou, Zi-Yi Dou, Jianwei Yang, Zhe Gan, Linjie Li, Chunyuan Li, Xiyang Dai, Harkirat Behl, Jianfeng Wang, Lu Yuan, et al. Generalized decoding for pixel, image, and language. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 151...
2023
-
[110]
3D Object Detection (3DOD)
Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Wang, Lijuan Wang, Jianfeng Gao, and Yong Jae Lee. Segment everything everywhere all at once. Advances in Neural Information Processing Systems, 36, 2024. 2, 3, 4, 7, 8 Mosaic3D: Foundation Dataset and Model for...
2024
-
[111]
In the first round, LLaV A-1.5 is prompted to generate an image caption describing the overall scene
-
[112]
In the second round, LLaV A-1.5 is prompted to extract entity names from the generated image caption
-
[113]
entity name A
In the final round, LLaV A-1.5 is prompted to generate de- tailed entity descriptions for each extracted entity name. During our implementation, we encountered inconsistencies in LLaV A-1.5’s response formats. To ensure structured and consistent entity-level text descriptions,...
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.