REVIEW 2 major objections 2 minor 105 references
Count Anything
T0 review · 2 major / 2 minor · reviewed 2026-06-28 · grok-4.3
Pith's one-line read Count Anything shows one model can count objects from text queries across scenes, microscopy, and remote sensing.
desk verdict Count Anything brings text guidance and dual counters to object counting with a new benchmark, but CLOC's construction needs verification to support the generalization claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Dual-granularity instance enumeration that pairs a Region-level Sparse Counter for large sparse targets with a Pixel-level Dense Counter for small dense targets, fused by Complementary Count Fusion.
What would settle it
Accuracy measurements on a held-out visual domain or object category absent from CLOC training that fall below the performance of existing domain-specific counters.
Extended reading notes
Core claim
Count Anything is a generalist model for text-guided object counting that replaces density maps with discrete instance points and performs dual-granularity enumeration: a Region-level Sparse Counter supplies anchors for large sparse targets while a Pixel-level Dense Counter predicts dense points for small crowded targets, with point-centric supervision and parameter-free Complementary Count Fusion enabling training on the heterogeneous CLOC collection that covers general scenes, remote sensing, histopathology, cellular microscopy, agriculture and microbiology.
Load-bearing premise
Reorganizing public datasets into CLOC produces a balanced benchmark that tests cross-domain generalization without annotation artifacts or domain leakage.
Editorial extensions
If this is right
- The same model outperforms prior open-world counting methods on accuracy and multi-domain generalization.
- Point-centric supervision lets the model learn from mixed annotation formats across source datasets.
- Output points supply both the count and the spatial locations of every detected instance.
- CLOC becomes a standard testbed for measuring whether counting models truly cross visual domains.
Reading between the lines
- Point-based counting may replace density maps in settings where instance locations are needed for downstream tasks.
- Natural-language prompts could let non-specialists request counts inside medical or satellite imagery without writing code.
- Adding text examples for a new domain might suffice for adaptation instead of collecting fresh labeled images.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces Count Anything, a text-guided object counting model that employs dual-granularity counters—a Region-level Sparse Counter for large sparse objects and a Pixel-level Dense Counter for small crowded ones—combined via parameter-free Complementary Count Fusion. It also presents CLOC, a new benchmark reorganizing existing datasets into six domains (General Scene, Remote Sensing, Histopathology, Cellular Microscopy, Agriculture, Microbiology) with ~220K images, 619 categories, and 15M instances, claiming superior multi-domain performance and generalization over existing open-world counting methods.
Significance. If the cross-domain results hold without benchmark artifacts, the work would provide a useful generalist baseline for text-conditioned counting, unifying previously fragmented domain-specific tasks. The point-centric supervision strategy that accommodates heterogeneous annotations (point, box, density) and the release of code are concrete strengths that support reproducibility and further research.
major comments (2)
- [CLOC construction / dataset section] CLOC construction (described in the abstract and the dataset section): the manuscript provides no analysis demonstrating absence of overlapping images/near-duplicates across the six domains, category-name collisions with inconsistent definitions, or annotation-style leakage (point vs. box vs. density maps) from the source collections. Because the headline claim of multi-domain generalization rests on CLOC being an unbiased test set, this omission is load-bearing.
- [Experiments section] Experiments section: the reported outperformance is presented without error bars, ablation tables isolating the contribution of each counter or the fusion step, or statistics on domain balance and category distribution within CLOC. This makes it impossible to assess whether the gains are robust or sensitive to post-hoc choices.
minor comments (2)
- [Abstract / Method] The abstract states the fusion is 'parameter-free' but does not explicitly define the fusion rule or show that no learned weights are involved; a short equation or pseudocode would clarify this.
- [Method] Notation for the two counters (Region-level Sparse Counter, Pixel-level Dense Counter) is introduced without a table comparing their supervision signals or output formats; a small comparison table would improve readability.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback and positive assessment of the work's potential significance. We address the two major comments below and commit to revisions that strengthen the manuscript without altering its core claims.
read point-by-point responses
-
Referee: [CLOC construction / dataset section] CLOC construction (described in the abstract and the dataset section): the manuscript provides no analysis demonstrating absence of overlapping images/near-duplicates across the six domains, category-name collisions with inconsistent definitions, or annotation-style leakage (point vs. box vs. density maps) from the source collections. Because the headline claim of multi-domain generalization rests on CLOC being an unbiased test set, this omission is load-bearing.
Authors: We agree that explicit verification of cross-domain image overlaps, near-duplicates, category-name consistency, and annotation-style leakage is necessary to support the multi-domain generalization claims. The current manuscript does not include such analysis. In the revision we will add a dedicated subsection reporting: (i) image-level deduplication checks (e.g., perceptual hashing and embedding similarity thresholds) across the six source collections, (ii) manual review of category-name collisions with harmonized definitions where needed, and (iii) confirmation that annotation formats were converted without leakage of supervision style into the evaluation splits. If any overlaps are found, we will document removal statistics. revision: yes
-
Referee: [Experiments section] Experiments section: the reported outperformance is presented without error bars, ablation tables isolating the contribution of each counter or the fusion step, or statistics on domain balance and category distribution within CLOC. This makes it impossible to assess whether the gains are robust or sensitive to post-hoc choices.
Authors: We concur that the absence of error bars, component-wise ablations, and dataset statistics limits interpretability of the reported gains. The revision will include: (i) standard deviations or confidence intervals over multiple random seeds for all main results, (ii) ablation tables that isolate the Region-level Sparse Counter, Pixel-level Dense Counter, and Complementary Count Fusion, and (iii) supplementary tables/figures showing per-domain image counts, category distributions, and instance-density histograms within CLOC. These additions will be placed in the Experiments section and appendix. revision: yes
Circularity Check
No significant circularity; claims rest on empirical evaluation of reorganized benchmark
full rationale
The paper's central claims rest on constructing CLOC by reorganizing existing public datasets and then reporting experimental accuracy of the proposed dual-granularity counters on that benchmark. No derivation step reduces a prediction or first-principles result to its own inputs by construction, no fitted parameter is relabeled as a prediction, and no load-bearing uniqueness theorem or ansatz is imported via self-citation. The performance numbers are obtained from standard train/test splits on the assembled data rather than being forced by the model definition itself.
Assumptions & free parameters
invented entities (2)
-
Region-level Sparse Counter
-
Pixel-level Dense Counter
Cite this review
Pith. "Pith review of Count Anything." pith.science (2026). https://pith.science/paper/U4XB4HSG
@misc{pith2026260530846,
author = {Pith},
title = {Pith review of: Count Anything},
year = {2026},
howpublished = {\url{https://pith.science/paper/U4XB4HSG}},
note = {Machine review of arXiv:2605.30846}
}
read the original abstract
Object counting remains fragmented across domain-specific datasets and task formulations, despite rapid progress in generalist vision models. Existing counting models are often tailored to scenarios such as crowds, vehicles, cells, crops, or remote-sensing objects, and thus struggle to generalize across categories, visual domains, object scales, and density distributions. In this paper, we study text-guided object counting across domains, where a model takes an image and a natural-language query as input and returns an instance-grounded set of target points whose cardinality gives the count. This formulation unifies category-conditioned counting with interpretable spatial localization. To support this setting, we construct CLOC, a Cross-domain Large-scale Object Counting dataset that reorganizes diverse public data sources into a unified benchmark. CLOC covers six visual domains: General Scene, Remote Sensing, Histopathology, Cellular Microscopy, Agriculture, and Microbiology, with about 220K images, 619 categories, and 15M object instances. Based on CLOC, we propose Count Anything, a generalist model for text-guided object counting. Unlike density-map-based methods, which dominate counting models, Count Anything adopts discrete instance points and performs dual-granularity instance enumeration. A Region-level Sparse Counter provides object-level anchors for large and sparse targets, while a Pixel-level Dense Counter handles small, crowded, and weakly bounded targets via dense point prediction. A point-centric supervision strategy enables learning from heterogeneous annotations, and Complementary Count Fusion combines both counters in a parameter-free manner. Extensive experiments show that Count Anything achieves strong accuracy and multi-domain generalization, outperforming existing open-world counting methods. Code is available at: https://github.com/Mengqi-Lei/count-anything.
Figures
Figures from the paper (14 more)
Reference graph
Works this paper leans on
-
[1]
Foundation models defining a new era in vision: A survey and outlook.IEEE Trans
Muhammad Awais, Muzammal Naseer, Salman Khan, Rao Muhammad Anwer, Hisham Cholakkal, Mubarak Shah, Ming-Hsuan Yang, and Fahad Shahbaz Khan. Foundation models defining a new era in vision: A survey and outlook.IEEE Trans. Pattern Anal. Mach. Intell., 47(4):2245–2264, 2025
2025
-
[2]
Vision foundation models in remote sensing: A survey.IEEE Geosci
Siqi Lu, Junlin Guo, James R Zimmer-Dauphinee, Jordan M Nieusma, Xiao Wang, Parker VanValkenburgh, Steven A Wernke, and Yuankai Huo. Vision foundation models in remote sensing: A survey.IEEE Geosci. Remote Sens. Mag., 13(3):190–215, 2025
2025
-
[3]
Florence: A New Foundation Model for Computer Vision
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. Florence: A new foundation model for computer vision.arXiv preprint arXiv:2111.11432, 2021
work page Pith review arXiv 2021
-
[4]
VLCounter: Text-aware visual representa- tion for zero-shot object counting
Seunggu Kang, WonJun Moon, Euiyeon Kim, and Jae-Pil Heo. VLCounter: Text-aware visual representa- tion for zero-shot object counting. InAAAI, volume 38, pages 2714–2722, 2024
2024
-
[5]
Zero-Shot Object Counting With Good Exemplars
Huilin Zhu, Jingling Yuan, Zhengwei Yang, Yu Guo, Zheng Wang, Xian Zhong, and Shengfeng He. Zero-Shot Object Counting With Good Exemplars. InEur . Conf. Comput. Vis., pages 368–385, 2024
2024
-
[6]
CountGD: Multi-modal open-world counting
Niki Amini-Naieni, Tengda Han, and Andrew Zisserman. CountGD: Multi-modal open-world counting. In Adv. Neural Inform. Process. Syst., volume 37, pages 48810–48837, 2024
2024
-
[7]
CountGD++: Generalized prompting for open-world counting
Niki Amini-Naieni and Andrew Zisserman. CountGD++: Generalized prompting for open-world counting. InIEEE Conf. Comput. Vis. Pattern Recog., 2026. 10
2026
-
[8]
Deep learning in crowd counting: A survey.CAAI Trans
Lijia Deng, Qinghua Zhou, Shuihua Wang, Juan Manuel Górriz, and Yudong Zhang. Deep learning in crowd counting: A survey.CAAI Trans. on Intell. Tech., 9(5):1043–1077, 2024
2024
Show all 105 references
-
[9]
Zero-shot object counting
Jingyi Xu, Hieu Le, Vu Nguyen, Viresh Ranjan, and Dimitris Samaras. Zero-shot object counting. InIEEE Conf. Comput. Vis. Pattern Recog., pages 15548–15557, 2023
2023
-
[10]
Rethinking counting and localization in crowds: A purely point-based framework
Qingyu Song, Changan Wang, Zhengkai Jiang, Yabiao Wang, Ying Tai, Chengjie Wang, Jilin Li, Feiyue Huang, and Yang Wu. Rethinking counting and localization in crowds: A purely point-based framework. InInt. Conf. Comput. Vis., pages 3365–3374, 2021
2021
-
[11]
Learning To Count Anything: Reference-less class-agnostic counting with weak supervision
Michael Hobley and Victor Prisacariu. Learning To Count Anything: Reference-less class-agnostic counting with weak supervision. InIEEE Conf. Comput. Vis. Pattern Recog. Worksh., 2023
2023
-
[12]
Learning To Count Everything
Viresh Ranjan, Udbhav Sharma, Thu Nguyen, and Minh Hoai. Learning To Count Everything. InIEEE Conf. Comput. Vis. Pattern Recog., pages 3394–3403, 2021
2021
-
[13]
Few-shot Object Counting and Detection
Thanh Nguyen, Chau Pham, Khoi Nguyen, and Minh Hoai. Few-shot Object Counting and Detection. In Eur . Conf. Comput. Vis., pages 348–365, 2022
2022
-
[14]
Single-image crowd counting via multi-column convolutional neural network
Yingying Zhang, Desen Zhou, Siqin Chen, Shenghua Gao, and Yi Ma. Single-image crowd counting via multi-column convolutional neural network. InIEEE Conf. Comput. Vis. Pattern Recog., pages 589–597, 2016
2016
-
[15]
Improving Point-Based Crowd Counting and Localization Based on Auxiliary Point Guidance
I-Hsiang Chen, Wei-Ting Chen, Yu-Wei Liu, Ming-Hsuan Yang, and Sy-Yen Kuo. Improving Point-Based Crowd Counting and Localization Based on Auxiliary Point Guidance. InEur . Conf. Comput. Vis., pages 428–444, 2024
2024
-
[16]
CCTrans: Simplifying and improving crowd counting with transformer.arXiv preprint arXiv:2109.14483, 2021
Ye Tian, Xiangxiang Chu, and Hongpeng Wang. CCTrans: Simplifying and improving crowd counting with transformer.arXiv preprint arXiv:2109.14483, 2021
2021
-
[17]
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft COCO: Common objects in context. InEur . Conf. Comput. Vis., pages 740–755, 2014
2014
-
[18]
Mark Everingham, Luc Van Gool, Christopher K. I. Williams, John Winn, and Andrew Zisserman. The PASCAL Visual Object Classes (VOC) Challenge.Int. J. Comput. Vis., 88(2):303–338, 2010
2010
-
[19]
Sindagi, Rajeev Yasarla, and Vishal M
Vishwanath A. Sindagi, Rajeev Yasarla, and Vishal M. Patel. JHU-CROWD++: Large-scale crowd counting dataset and a benchmark method.IEEE Trans. Pattern Anal. Mach. Intell., 44(5):2594–2609, 2022
2022
-
[20]
NuInsSeg: A fully annotated dataset for nuclei instance segmentation in H&E-stained histological images.Sci
Amirreza Mahbod, Christine Polak, Katharina Feldmann, Rumsha Khan, Katharina Gelles, Georg Dorffner, Ramona Woitek, Sepideh Hatamikia, and Isabella Ellinger. NuInsSeg: A fully annotated dataset for nuclei instance segmentation in H&E-stained histological images.Sci. Data, 11(1...
2024
-
[21]
A Dataset and a Technique for Generalized Nuclear Segmentation for Computational Pathology.IEEE Trans
Neeraj Kumar, Ruchika Verma, Sanuj Sharma, Surabhi Bhargava, Abhishek Vahadane, and Amit Sethi. A Dataset and a Technique for Generalized Nuclear Segmentation for Computational Pathology.IEEE Trans. Med. Imag., 36(7):1550–1560, 2017
2017
-
[22]
CellBinDB: A large-scale multimodal annotated dataset for cell segmentation with benchmarking of universal models.GigaScience, 14:giaf069, 2025
Can Shi, Jinghong Fan, Zhonghan Deng, Huanlin Liu, Qiang Kang, Yumei Li, Jing Guo, Jingwen Wang, Jinjiang Gong, Sha Liao, Ao Chen, Ying Zhang, and Mei Li. CellBinDB: A large-scale multimodal annotated dataset for cell segmentation with benchmarking of universal models.GigaScie...
2025
-
[23]
Objects365: A large-scale, high-quality dataset for object detection
Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. Objects365: A large-scale, high-quality dataset for object detection. InInt. Conf. Comput. Vis., pages 8430–8439, 2019
2019
-
[24]
Detection and Tracking Meet Drones Challenge.IEEE Trans
Pengfei Zhu, Longyin Wen, Dawei Du, Xiao Bian, Heng Fan, Qinghua Hu, and Haibin Ling. Detection and Tracking Meet Drones Challenge.IEEE Trans. Pattern Anal. Mach. Intell., 44(11):7380–7399, 2022
2022
-
[25]
Ahmed Raza, Nasir M
Ruchika Verma, Neeraj Kumar, Abhijeet Patil, Nikhil Cherian Kurian, Swapnil Rane, Simon Graham, Quoc Dang Vu, Mieke Zwager, Shan E. Ahmed Raza, Nasir M. Rajpoot, et al. MoNuSAC2020: A multi-organ nuclei segmentation and classification challenge.IEEE Trans. Med. Imag., 40(12):3...
2021
-
[26]
Venkatesh Babu
Deepak Babu Sam, Skand Vishwanath Peri, Mukuntha Narayanan Sundararaman, Amogh Kamath, and R. Venkatesh Babu. Locate, size, and count: Accurately resolving people in dense crowds via detection. IEEE Trans. Pattern Anal. Mach. Intell., 43(8):2739–2751, 2021. 11
2021
-
[27]
Meng-Ru Hsieh, Yen-Liang Lin, and Winston H. Hsu. Drone-Based Object Counting by Spatially Regularized Regional Proposal Networks. InInt. Conf. Comput. Vis., pages 4145–4153, 2017
2017
-
[28]
Distribution Matching for Crowd Counting
Boyu Wang, Huidong Liu, Dimitris Samaras, and Minh Hoai Nguyen. Distribution Matching for Crowd Counting. InAdv. Neural Inform. Process. Syst., pages 1595–1607, 2020
2020
-
[29]
CLIP-Count: Towards text-guided zero-shot object counting
Ruixiang Jiang, Lingbo Liu, and Changwen Chen. CLIP-Count: Towards text-guided zero-shot object counting. InACM Int. Conf. Multimedia, pages 4535–4545, 2023
2023
-
[30]
Consistency-aware anchor pyramid network for crowd localization.IEEE Trans
Xinyan Liu, Guorong Li, Yuankai Qi, Zhenjun Han, Anton van den Hengel, Nicu Sebe, Ming-Hsuan Yang, and Qingming Huang. Consistency-aware anchor pyramid network for crowd localization.IEEE Trans. Pattern Anal. Mach. Intell., 2024
2024
-
[31]
Kevin Zhou, and Jie Chen
Zhongyi Huang, Yao Ding, Guoli Song, Lin Wang, Ruizhe Geng, Hongliang He, Shan Du, Xia Liu, Yonghong Tian, Yongsheng Liang, S. Kevin Zhou, and Jie Chen. BCData: A large-scale dataset and benchmark for cell detection and counting. InMed. Image Comput. Comput. Assist. Interv., p...
2020
-
[32]
TasselNet: Counting maize tassels in the wild via local counts regression network.Plant Methods, 13(1):79, 2017
Hao Lu, Zhiguo Cao, Yang Xiao, Bohan Zhuang, and Chunhua Shen. TasselNet: Counting maize tassels in the wild via local counts regression network.Plant Methods, 13(1):79, 2017
2017
-
[33]
LVIS: A dataset for large vocabulary instance segmentation
Agrim Gupta, Piotr Dollár, and Ross Girshick. LVIS: A dataset for large vocabulary instance segmentation. InIEEE Conf. Comput. Vis. Pattern Recog., pages 5356–5364, 2019
2019
-
[34]
Plummer, Liwei Wang, Chris M
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. Flickr30k Entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. InInt. Conf. Comput. Vis., pages 2641–2649, 2015
2015
-
[35]
SAM 3: Segment anything with concepts
Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Didac Suris, Chaitanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, et al. SAM 3: Segment anything with concepts. InInt. Conf. Learn. Represent., 2026
2026
-
[36]
End-to-End Object Detection with Transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-End Object Detection with Transformers. InEur . Conf. Comput. Vis., pages 213–229, 2020
2020
-
[37]
A Low-Shot Object Counting Network With Iterative Prototype Adaptation
Nikola Ðuki´c, Alan Lukežiˇc, Vitjan Zavrtanik, and Matej Kristan. A Low-Shot Object Counting Network With Iterative Prototype Adaptation. InInt. Conf. Comput. Vis., pages 18872–18881, 2023
2023
-
[38]
Open-World Text-Specified Object Counting
Niki Amini-Naieni, Kiana Amini-Naieni, Tengda Han, and Andrew Zisserman. Open-World Text-Specified Object Counting. InBrit. Mach. Vis. Conf., 2023
2023
-
[39]
YOLO-Count: Differentiable object counting for text-to-image generation
Guanning Zeng, Xiang Zhang, Zirui Wang, Haiyang Xu, Zeyuan Chen, Bingnan Li, and Zhuowen Tu. YOLO-Count: Differentiable object counting for text-to-image generation. InInt. Conf. Comput. Vis., pages 16765–16775, 2025
2025
-
[40]
Yifei Qian, Zhongliang Guo, Bowen Deng, Chun Tong Lei, Shuai Zhao, Chun Pong Lau, Xiaopeng Hong, and Michael P. Pound. T2ICount: Enhancing cross-modal understanding for zero-shot counting. InIEEE Conf. Comput. Vis. Pattern Recog., pages 25336–25345, 2025
2025
-
[41]
CountSE: Soft exemplar open-set object counting
Shuai Liu, Peng Zhang, Shiwei Zhang, and Wei Ke. CountSE: Soft exemplar open-set object counting. In Int. Conf. Comput. Vis., pages 21536–21546, 2025
2025
-
[42]
Grounded Language-Image Pre-Training
Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, et al. Grounded Language-Image Pre-Training. InIEEE Conf. Comput. Vis. Pattern Recog., pages 10965–10975, 2022
2022
-
[43]
YOLO-World: Real- time open-vocabulary object detection
Tianheng Cheng, Lin Song, Yixiao Ge, Wenyu Liu, Xinggang Wang, and Ying Shan. YOLO-World: Real- time open-vocabulary object detection. InIEEE Conf. Comput. Vis. Pattern Recog., pages 16901–16911, 2024
2024
-
[44]
Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, and Lei Zhang. Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection. InEur . Conf. Comput. Vis., pages 38–55, 2024
2024
-
[45]
YOLOE: Real-time seeing anything
Ao Wang, Lihao Liu, Hui Chen, Zijia Lin, Jungong Han, and Guiguang Ding. YOLOE: Real-time seeing anything. InInt. Conf. Comput. Vis., pages 24591–24602, 2025. 12
2025
-
[46]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. InInt. Conf. Learn. Represent., 2022
2022
-
[47]
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. InInt. Conf. Learn. Represent., 2019
2019
-
[48]
Redesigning Multi-Scale Neural Network for Crowd Counting.IEEE Trans
Zhipeng Du, Miaojing Shi, Jiankang Deng, and Stefanos Zafeiriou. Redesigning Multi-Scale Neural Network for Crowd Counting.IEEE Trans. Image Process., 32:3664–3678, 2023
2023
-
[49]
CSRNet: Dilated convolutional neural networks for understanding the highly congested scenes
Yuhong Li, Xiaofan Zhang, and Deming Chen. CSRNet: Dilated convolutional neural networks for understanding the highly congested scenes. InIEEE Conf. Comput. Vis. Pattern Recog., pages 1091–1100, 2018
2018
-
[50]
Scale Aggregation Network for Accurate and Efficient Crowd Counting
Xinkun Cao, Zhipeng Wang, Yanyun Zhao, and Fei Su. Scale Aggregation Network for Accurate and Efficient Crowd Counting. InEur . Conf. Comput. Vis., pages 734–750, 2018
2018
-
[51]
Bayesian loss for crowd count estimation with point supervision
Zhiheng Ma, Xing Wei, Xiaopeng Hong, and Yihong Gong. Bayesian loss for crowd count estimation with point supervision. InInt. Conf. Comput. Vis., pages 6142–6151, 2019
2019
-
[52]
CounTR: Transformer-based generalised visual counting
Chang Liu, Yujie Zhong, Andrew Zisserman, and Weidi Xie. CounTR: Transformer-based generalised visual counting. InBrit. Mach. Vis. Conf., 2022
2022
-
[53]
Point, Segment and Count: A generalized framework for object counting
Zhizhong Huang, Mingliang Dai, Yi Zhang, Junping Zhang, and Hongming Shan. Point, Segment and Count: A generalized framework for object counting. InIEEE Conf. Comput. Vis. Pattern Recog., pages 17067–17076, 2024
2024
-
[54]
DA VE: A detect-and-verify paradigm for low-shot counting
Jer Pelhan, Alan Lukežiˇc, Vitjan Zavrtanik, and Matej Kristan. DA VE: A detect-and-verify paradigm for low-shot counting. InIEEE Conf. Comput. Vis. Pattern Recog., pages 23293–23302, 2024
2024
-
[55]
Open-World Object Counting in Videos
Niki Amini-Naieni and Andrew Zisserman. Open-World Object Counting in Videos. InAAAI, volume 40, pages 2300–2308, 2026
2026
-
[56]
NWPU-Crowd: A large-scale benchmark for crowd counting and localization.IEEE Trans
Qi Wang, Junyu Gao, Wei Lin, and Xuelong Li. NWPU-Crowd: A large-scale benchmark for crowd counting and localization.IEEE Trans. Pattern Anal. Mach. Intell., 43(6):2141–2149, 2021
2021
-
[57]
Badhon, Curtis Pozniak, Benoit de Solan, Andreas Hund, Scott C
Etienne David, Simon Madec, Pouria Sadeghi-Tehran, Helge Aasen, Bangyou Zheng, Shouyang Liu, Norbert Kirchgessner, Goro Ishikawa, Koichi Nagasawa, Minhajul A. Badhon, Curtis Pozniak, Benoit de Solan, Andreas Hund, Scott C. Chapman, Frédéric Baret, Ian Stavness, and Wei Guo. Gl...
2020
-
[58]
AGAR: A microbial colony dataset for deep learning detection, 2021
Sylwia Majchrowska, Jarosław Pawłowski, Grzegorz Guła, Tomasz Bonus, Agata Hanas, Adam Loch, Agnieszka Pawlak, Justyna Roszkowiak, Tomasz Golan, and Zuzanna Drulis-Kawa. AGAR: A microbial colony dataset for deep learning detection, 2021
2021
-
[59]
OmniCount: Multi-label object counting with semantic-geometric priors
Anindya Mondal, Sauradip Nag, Xiatian Zhu, and Anjan Dutta. OmniCount: Multi-label object counting with semantic-geometric priors. InAAAI, volume 39, pages 19537–19545, 2025
2025
-
[60]
DOTA: A large-scale dataset for object detection in aerial images
Gui-Song Xia, Xiang Bai, Jian Ding, Zhen Zhu, Serge Belongie, Jiebo Luo, Mihai Datcu, Marcello Pelillo, and Liangpei Zhang. DOTA: A large-scale dataset for object detection in aerial images. InIEEE Conf. Comput. Vis. Pattern Recog., pages 3974–3983, 2018
2018
-
[61]
xView: Objects in context in overhead imagery.arXiv preprint arXiv:1802.07856, 2018
Darius Lam, Richard Kuzma, Kevin McGee, Samuel Dooley, Michael Laielli, Matthew Klaric, Yaroslav Bulatov, and Brendan McCord. xView: Objects in context in overhead imagery.arXiv preprint arXiv:1802.07856, 2018
2018 arXiv
-
[62]
The Cityscapes Dataset for Semantic Urban Scene Understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benen- son, Uwe Franke, Stefan Roth, and Bernt Schiele. The Cityscapes Dataset for Semantic Urban Scene Understanding. InIEEE Conf. Comput. Vis. Pattern Recog., pages 3213–3223, 2016
2016
-
[63]
Jackson, Nabeel Khalid, Nicola Bevan, Timothy Dale, Andreas Dengel, Sheraz Ahmed, Johan Trygg, and Rickard Sjögren
Christoffer Edlund, Timothy R. Jackson, Nabeel Khalid, Nicola Bevan, Timothy Dale, Andreas Dengel, Sheraz Ahmed, Johan Trygg, and Rickard Sjögren. LIVECell: A large-scale dataset for label-free live cell segmentation.Nat. Methods, 18(9):1038–1045, 2021
2021
-
[64]
SoybeanNet: Transformer-based convolutional neural network for soybean pod counting from unmanned aerial vehicle (UA V) images.Comput
Jiajia Li, Raju Thada Magar, Dong Chen, Feng Lin, Dechun Wang, Xiang Yin, Weichao Zhuang, and Zhaojian Li. SoybeanNet: Transformer-based convolutional neural network for soybean pod counting from unmanned aerial vehicle (UA V) images.Comput. Electron. Agric., 220:108861, 2024. 13
2024
-
[65]
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning Transferable Visual Models From Natural Language Supervision. InInt. Conf. Mach. L...
2021
-
[66]
Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick. Segment Anything. InInt. Conf. Comput. Vis., pages 4015–4026, 2023
2023
-
[67]
MedMNIST v2: A large-scale lightweight benchmark for 2D and 3D biomedical image classification
Jiancheng Yang, Rui Shi, Donglai Wei, Zequan Liu, Lin Zhao, Bilian Ke, Hanspeter Pfister, and Bingbing Ni. MedMNIST v2: A large-scale lightweight benchmark for 2D and 3D biomedical image classification. Sci. Data, 10(1):41, 2023
2023
-
[68]
Landman, Geert Litjens, Bjoern Menze, Olaf Ronneberger, Ronald M
Michela Antonelli, Annika Reinke, Spyridon Bakas, Keyvan Farahani, Annette Kopp-Schneider, Bennett A. Landman, Geert Litjens, Bjoern Menze, Olaf Ronneberger, Ronald M. Summers, Bram van Ginneken, et al. The Medical Segmentation Decathlon.Nat. Commun., 13(1):4128, 2022
2022
-
[69]
MSeg: A composite dataset for multi-domain semantic segmentation
John Lambert, Zhuang Liu, Ozan Sener, James Hays, and Vladlen Koltun. MSeg: A composite dataset for multi-domain semantic segmentation. InIEEE Conf. Comput. Vis. Pattern Recog., pages 2879–2888, 2020
2020
-
[70]
BigDetection: A large-scale benchmark for improved object detector pre-training
Likun Cai, Zhi Zhang, Yi Zhu, Li Zhang, Mu Li, and Xiangyang Xue. BigDetection: A large-scale benchmark for improved object detector pre-training. InIEEE Conf. Comput. Vis. Pattern Recog. Worksh., pages 4777–4787, 2022
2022
-
[71]
rebar counting dataset
fyp2. rebar counting dataset. https://universe.roboflow.com/fyp2-czq30/ rebar-counting-12vha, September 2024. visited on 2026-05-02
2024
-
[72]
Paulo R. L. de Almeida, Luiz S. Oliveira, Alceu S. Britto Jr., Eunelson J. Silva Jr., and Alessandro L. Koerich. PKLot: A robust dataset for parking lot classification.Expert Syst. Appl., 42(11):4937–4949, 2015
2015
-
[73]
Object Detection in Optical Remote Sensing Images: A survey and a new benchmark.ISPRS J
Ke Li, Gang Wan, Gong Cheng, Liqiu Meng, and Junwei Han. Object Detection in Optical Remote Sensing Images: A survey and a new benchmark.ISPRS J. Photogramm. Remote Sens., 159:296–307, 2020
2020
-
[74]
Object Detection in Aerial Images: A large-scale benchmark and challenges.IEEE Trans
Jian Ding, Nan Xue, Gui-Song Xia, Xiang Bai, Wen Yang, Michael Ying Yang, Serge Belongie, Jiebo Luo, Mihai Datcu, Marcello Pelillo, and Liangpei Zhang. Object Detection in Aerial Images: A large-scale benchmark and challenges.IEEE Trans. Pattern Anal. Mach. Intell., 44(11):777...
2022
-
[75]
Learning RoI Transformer for Oriented Object Detection in Aerial Images
Jian Ding, Nan Xue, Yang Long, Gui-Song Xia, and Qikai Lu. Learning RoI Transformer for Oriented Object Detection in Aerial Images. InIEEE Conf. Comput. Vis. Pattern Recog., pages 2849–2858, 2019
2019
-
[76]
Detection, Tracking, and Counting Meets Drones in Crowds: A benchmark
Longyin Wen, Dawei Du, Pengfei Zhu, Qinghua Hu, Qilong Wang, Liefeng Bo, and Siwei Lyu. Detection, Tracking, and Counting Meets Drones in Crowds: A benchmark. InIEEE Conf. Comput. Vis. Pattern Recog., pages 7812–7821, 2021
2021
-
[77]
NWPU-MOC: A benchmark for fine-grained multicategory object counting in aerial images.IEEE Trans
Junyu Gao, Liangliang Zhao, and Xuelong Li. NWPU-MOC: A benchmark for fine-grained multicategory object counting in aerial images.IEEE Trans. Geosci. Remote Sens., 62:1–14, 2024
2024
-
[78]
Multi-Class Geospatial Object Detection and Geographic Image Classification Based on Collection of Part Detectors.ISPRS J
Gong Cheng, Junwei Han, Peicheng Zhou, and Lei Guo. Multi-Class Geospatial Object Detection and Geographic Image Classification Based on Collection of Part Detectors.ISPRS J. Photogramm. Remote Sens., 98:119–132, 2014
2014
-
[79]
A Survey on Object Detection in Optical Remote Sensing Images.ISPRS J
Gong Cheng and Junwei Han. A Survey on Object Detection in Optical Remote Sensing Images.ISPRS J. Photogramm. Remote Sens., 117:11–28, 2016
2016
-
[80]
Learning Rotation-Invariant Convolutional Neural Networks for Object Detection in VHR Optical Remote Sensing Images.IEEE Trans
Gong Cheng, Peicheng Zhou, and Junwei Han. Learning Rotation-Invariant Convolutional Neural Networks for Object Detection in VHR Optical Remote Sensing Images.IEEE Trans. Geosci. Remote Sens., 54(12):7405–7415, 2016
2016
-
[81]
Accurate Object Localization in Remote Sensing Images Based on Convolutional Neural Networks.IEEE Trans
Yang Long, Yiping Gong, Zhifeng Xiao, and Qing Liu. Accurate Object Localization in Remote Sensing Images Based on Convolutional Neural Networks.IEEE Trans. Geosci. Remote Sens., 55(5):2486–2498, 2017
2017
-
[82]
Elliptic Fourier Transformation-Based Histograms of Oriented Gradients for Rotationally Invariant Object Detection in Remote-Sensing Images.Int
Zhifeng Xiao, Qing Liu, Gefu Tang, and Xiaofang Zhai. Elliptic Fourier Transformation-Based Histograms of Oriented Gradients for Rotationally Invariant Object Detection in Remote-Sensing Images.Int. J. Remote Sens., 36(2):618–644, 2015
2015
-
[83]
Enhancing People Localisation in Drone Imagery for Better Crowd Management by Utilising Every Pixel in High-Resolution Images.arXiv preprint arXiv:2502.04014, 2025
Bartosz Ptak and Marek Kraft. Enhancing People Localisation in Drone Imagery for Better Crowd Management by Utilising Every Pixel in High-Resolution Images.arXiv preprint arXiv:2502.04014, 2025. 14
2025
-
[84]
CoNIC Challenge: Pushing the frontiers of nuclear detection, segmentation, classification and counting.Med
Simon Graham, Quoc Dang Vu, Mostafa Jahanifar, Martin Weigert, Uwe Schmidt, Wenhua Zhang, Jun Zhang, Sen Yang, Jinxi Xiang, Xiyue Wang, et al. CoNIC Challenge: Pushing the frontiers of nuclear detection, segmentation, classification and counting.Med. Image Anal., 92:103047, 2024
2024
-
[85]
EndoNuke: Nuclei detection dataset for estrogen and progesterone stained IHC endometrium scans.Data, 7(6):75, 2022
Anton Naumov, Egor Ushakov, Andrey Ivanov, Konstantin Midiber, Tatyana Khovanskaya, Alexandra Konyukova, Polina Vishnyakova, Sergei Nora, Liudmila Mikhaleva, Timur Fatkhudinov, and Evgeny Karpulevich. EndoNuke: Nuclei detection dataset for estrogen and progesterone stained IHC...
2022
-
[86]
Lizard: A large-scale dataset for colonic nuclear instance segmentation and classification
Simon Graham, Mostafa Jahanifar, Ayesha Azam, Mohammed Nimir, Yee-Wah Tsang, Katherine Dodd, Emily Hero, Harvir Sahota, Atisha Tank, Ksenija Benes, et al. Lizard: A large-scale dataset for colonic nuclear instance segmentation and classification. InInt. Conf. Comput. Vis. Work...
2021
-
[87]
A Multi-Organ Nucleus Segmentation Challenge.IEEE Trans
Neeraj Kumar, Ruchika Verma, Deepak Anand, Yanning Zhou, Omer Fahri Onder, Efstratios Tsougenis, Hao Chen, Pheng-Ann Heng, Jiahui Li, Zhiqiang Hu, et al. A Multi-Organ Nucleus Segmentation Challenge.IEEE Trans. Med. Imag., 39(5):1380–1391, 2020
2020
-
[88]
Atteya, Hagar Hussein, Kareem Hosny Mohammed, Ehab Hafiz, Maha A
Mohamed Amgad, Lamees A. Atteya, Hagar Hussein, Kareem Hosny Mohammed, Ehab Hafiz, Maha A. T. Elsebaie, Ahmed M. Alhusseiny, Mohamed Atef AlMoslemany, Abdelmagid M. Elmatboly, Philip A. Pappalardo, Rokia Adel Sakr, et al. NuCLS: A scalable crowdsourcing approach and dataset fo...
2022
-
[89]
Lambert, and Bachir El Debs
Mathieu Gendarme, Annika M. Lambert, and Bachir El Debs. BriFiSeg: A deep learning-based method for semantic and instance segmentation of nuclei in brightfield images.arXiv preprint arXiv:2211.03072, 2022
2022
-
[90]
Learning To Count Objects in Images
Victor Lempitsky and Andrew Zisserman. Learning To Count Objects in Images. InAdv. Neural Inform. Process. Syst., pages 1324–1332, 2010
2010
-
[91]
Laine, Pedro M
Christoph Spahn, Estibaliz Gómez-de Mariscal, Romain F. Laine, Pedro M. Pereira, Lucas von Chamier, Mia Conduit, Mariana G. Pinho, Guillaume Jacquemet, Séamus Holden, Mike Heilemann, et al. DeepBacs for multi-task bacterial image analysis using open-source deep learning approa...
2022
-
[92]
Etienne David, Mario Serouart, Daniel Smith, Simon Madec, Kaaviya Velumani, Shouyang Liu, Xu Wang, Francisco Pinto, Shahameh Shafiee, Izzat S. A. Tahir, et al. Global Wheat Head Detection 2021: An improved dataset for benchmarking wheat head detection methods.Plant Phenomics, ...
2021
-
[93]
people’s heads
Luca Ciampi, Ali Azmoudeh, Elif Ecem Akbaba, Erdi Sarıta¸ s, Ziya Ata Yazıcı, Hazım Kemal Ekenel, Giuseppe Amato, and Fabrizio Falchi. A Survey on Class-Agnostic Counting: Advancements from reference-based to open-world text-guided approaches.Comput. Vis. Image Underst., 267:1...
2026
-
[94]
Some source datasets adopt standard COCO-style JSON annotations, such as Objects365-2020, FSCD-LVIS, and LIVECell
Parsing JSON-based annotation protocols.The original datasets vary substantially in their annotation protocols. Some source datasets adopt standard COCO-style JSON annotations, such as Objects365-2020, FSCD-LVIS, and LIVECell. These annotations are typically organized around i...
2020
-
[95]
VOC-style datasets use XML files to record the target category, bounding box, difficult flag, and related fields for each image
Parsing non-JSON annotation protocols.Beyond JSON annotations, we also process a variety of non-JSON annotation protocols. VOC-style datasets use XML files to record the target category, bounding box, difficult flag, and related fields for each image. Medical datasets such as ...
-
[96]
35439": {
Unified instance representation.During conversion, different types of instance annotations are unified into two basic representations: point and bbox. For instances that already provide bounding boxes, we retain their axis-aligned bounding boxes and generate corresponding coun...
-
[97]
Public datasets often adopt different naming conventions, such as differences in capitalization, singular/plural forms, spaces, underscores, and special characters
Category name normalization.We normalize category names so that categories from different data sources use consistent written forms. Public datasets often adopt different naming conventions, such as differences in capitalization, singular/plural forms, spaces, underscores, and...
-
[98]
one instance, one count
Removing categories unsuitable for counting.We further remove categories that are not suitable as instance-level counting targets. Not all original categories in multi-source datasets can be stably associated with the “one instance, one count” assumption. Some categories have ...
-
[99]
For category consolidation in a multi-source dataset, it is not appropriate to simply promote all subclasses to their parent class
Semantic category group construction.We construct semantic category groups for some fine- grained categories to support the later seen / unseen split. For category consolidation in a multi-source dataset, it is not appropriate to simply promote all subclasses to their parent c...
-
[100]
Since the dataset covers multiple visual domains, image resolutions vary substantially across data sources
High-resolution image normalization.After category consolidation, some extremely high- resolution images need to be normalized in scale. Since the dataset covers multiple visual domains, image resolutions vary substantially across data sources. For example, DOTA [60, 74, 75], ...
-
[101]
Cropped sample generation.Cropped samples are generated by extracting local target regions from original images. Unlike simple random cropping, our cropping process prioritizes local windows whose target counts fall into predefined count ranges, thereby supplementing medium- a...
-
[102]
Stitched sample generation.Stitched sample generation also follows a target-count-oriented strategy. We first select candidate image-patch combinations according to predefined target-count ranges, so that the stitched samples can supplement specified medium- and high-count int...
-
[103]
In other words, cropped and stitched samples are retained only in the training set, while the validation and test sets do not contain any samples generated by cropping or stitching
Derived-sample isolation.To avoid evaluation leakage introduced by data augmentation, we first construct the training, validation, and test sets at the original-sample level, and then generate cropped and stitched samples only from original images in the training set. In other...
-
[104]
To ensure stable cross-domain evaluation, the training, validation, and test sets all cover every visual domain
Visual-domain proportion constraint.The split also explicitly considers the distribution of the six visual domains. To ensure stable cross-domain evaluation, the training, validation, and test sets all cover every visual domain. Under the constraints of derived-sample isolatio...
-
[105]
Existing category-conditioned counting datasets usually construct 39 unseen-category evaluation by ensuring that category names do not overlap
Strict seen / unseen splitting based on category groups.For visual domains with rich category spaces, such as General Scene and Remote Sensing, we construct a seen / unseen evaluation setting to assess category generalization. Existing category-conditioned counting datasets us...
Reviewed June 28, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.