REVIEW 3 major objections 4 minor 92 references
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper introduces ODOR, a dataset of 38,116 object-level annotations across 4,712 artworks and 139 fine-grained categories, built to make object detection face small, dense, off-center, and smell-related objects in paintings.
desk verdict A genuinely useful new benchmark for artwork object detection, but its fine-grained label quality is unmeasured and needs to be demonstrated before the difficulty claims are taken at face value. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the dataset's two-level annotation hierarchy combined with its acquisition and analysis protocol. All 139 fine-grained classes sit under supercategories chosen for visual and olfactory similarity rather than biological taxonomy, so a detector can fall back to a coarse prediction when species-level confidence is low; this hierarchy is what makes the supercategory-evaluation experiment meaningful. The authors also use standard COCO-format annotations, a crowd-annotation pipeline with multiple rounds of expert art-history correction, and statistical characterizations of spatial distribution, occlusion, and box-size histograms to establish the dataset's difficulty. On top of this, a benchmark of five detector families (two-stage, transformer-based, and one-stage) with varying backbones provides the performance numbers that make the challenge concrete.
What would settle it
Independently re-annotate a random sample of about 200 ODOR test images with two art-history experts and measure per-class agreement on the fine-grained categories; low agreement on classes such as flower species or drinking vessels would mean the class-level AP baselines partly measure annotation noise rather than detector ability. In parallel, a chi-square goodness-of-fit test comparing object-centre positions to a uniform distribution over the canvas would settle whether the claimed full-canvas spread is real or an artifact of visual inspection.
Extended reading notes
Core claim
The central discovery is the dataset itself. ODOR contains 38,116 object-level bounding-box annotations in 4,712 historical artworks, organized into 139 fine-grained classes under eight pragmatic supercategories (flower, fruit, vegetable, mammal, drinking vessel, smoking equipment, smoke-related, and other). Its distinctive statistical properties are dense and overlapping instances (on average 8.1 boxes and 3.7 classes per image), a long-tailed class distribution in which some classes have only 16 training instances, a high share of very small objects covering less than 4 percent of image area, and object centres spread across the whole canvas instead of clustered in the middle. The paper demonstrates that these properties make detection genuinely hard: the best evaluated configuration, a DINO detector with a FocalNet-L backbone, reaches 22.6 AP; relaxing the task to supercategories raises AP substantially, showing that fine-grained distinctions are a main source of difficulty. The authors position ODOR as the second-largest public artwork detection dataset by image count and the most complex in terms of class count and annotation density.
Load-bearing premise
The whole benchmark rests on the accuracy of the 38,116 fine-grained labels, which were crowd-annotated and expert-corrected but with no reported inter-annotator agreement, spot-check error rate, or formal QA protocol, so every AP number inherits any unmeasured label noise.
Editorial extensions
If this is right
- The best baseline reaching only 22.6 AP shows that current detectors leave a wide margin on dense, fine-grained, small-object artwork detection.
- The 35 to 60 percent gain from evaluating at supercategory level means fine-grained class boundaries, not localisation, are the dominant source of difficulty for the strongest models.
- Detection-specific pretraining on natural images appears to transfer well: a COCO-pretrained one-stage detector roughly matches much heavier transformer backbones, especially on small objects.
- The dataset enables quantitative art-history studies of olfactory references, such as tracing the prevalence of roses in still-life paintings over centuries.
- Researchers needing robustness to off-centre, occluded, and small objects can use ODOR as a complement to natural-image benchmarks such as COCO.
Reading between the lines
- Because label noise is unreported, a practical next step is measuring inter-annotator agreement on the fine-grained classes; if it is low, category-level AP rankings may need to be re-read as rankings of annotation ambiguity.
- The supercategory hierarchy invites hierarchical or open-vocabulary detectors: a model that predicts the coarse class with calibrated uncertainty could be more useful for art historians than one forced to guess among 139 species-level labels.
- The dataset's class distribution reflects olfactory relevance, not general art-historical salience, so detectors trained on it may not transfer to arbitrary artwork-detection tasks without reweighting or fine-tuning.
- Because several images are only available through external source links, benchmark reproducibility depends on link persistence; hashed copies or a complete mirror would make future comparisons more robust.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces ODOR, a dataset of 4,712 artwork images with 38,116 object-level bounding-box annotations across 139 fine-grained categories and 8 supercategories, collected in the Odeuropa project with a focus on olfactory references. The authors report statistical analyses of category distribution, spatial spread, occlusion, and object sizes, and provide extensive baselines for five detector families (F-RCNN, DINO, YOLO-v8, FCOS, MADet), including supercategory-wise and class-wise evaluations. The paper claims that ODOR is the second-largest public artwork detection dataset by image count and the most complex by class count and annotation density, and that it offers a challenging benchmark for small, occluded, and peripheral objects in artworks.
Significance. If the label-quality concerns are resolved, ODOR would be a valuable and distinctive public resource for the computational humanities: it is the only artwork detection dataset focused on fine-grained olfactory-relevant categories, it exhibits a genuinely long-tailed distribution and dense, overlapping instances, and it is released with metadata, download tools, and reproducible baseline code. The baseline study is a concrete strength: the authors evaluate multiple modern detectors, report size- and supercategory-stratified AP, and make code and data publicly available (Zenodo, GitHub, HuggingFace), which lowers the barrier for follow-up work. The dataset has already been used in the ICPR 2022 ODOR challenge, indicating real community uptake. The central claims about dataset complexity and difficulty are, however, currently weakened by the absence of quantitative label-quality evidence and by at least one inaccurate comparative statement.
major comments (3)
- [§3.2, §5] The annotation process is described only qualitatively: Section 3.2 states that AMT annotations were 'manually checked by experts' with 'multiple rounds of corrections', and Section 5 concedes that 'an inherent level of ambiguity persists.' No inter-annotator agreement, spot-check error rate, per-class confusion statistics, or adjudication protocol is reported. This is load-bearing because the dataset's headline properties—139 fine-grained categories, subtle distinctions among flower species and drinking vessels—and all downstream quantities (baseline AP in Tables 4–6, class-wise AP and Spearman correlations in Table 7, and the supercategory gains in Table 6) inherit the unmeasured label noise. I request a quantitative QA component, such as Cohen's kappa on a held-out subset, a per-class error analysis on a random sample, or a sensitivity analysis showing that the main conclusions are robust to plausible label-error rates.
- [§6, Table 1] The conclusion states that ODOR is 'the second-largest public artwork dataset with object-level annotations in terms of the number of images.' This is contradicted by the paper's own Table 1, which lists DEArt with 15,157 images and Human-Art with 50,000 images—both larger than ODOR's 4,712. The claim must be corrected or explicitly qualified (e.g., excluding person-only or semi-automatically annotated datasets), otherwise the central positioning of the dataset is overstated.
- [§3.4, Fig. 7] The 'spreaded' property—one of the three characteristics advertised in the title—is supported only by visual comparison of heatmaps in Fig. 7. A quantitative test comparing the normalized object-center distributions of ODOR with PeopleArt, IconArt, and DEArt (e.g., a chi-square or Kolmogorov–Smirnov test) would make the claim falsifiable and is necessary to substantiate the full-canvas distribution assertion, which the paper explicitly contrasts with the center-biased distributions of other datasets.
minor comments (4)
- [Title] The title has a spacing typo: 'TheObject Detection' should be 'The Object Detection.'
- [§3.1] The sentence 'On average they have a width of 653 and a width of 641 pixels' should presumably read 'a width of 653 and a height of 641 pixels.'
- [§4.1] The phrase 'Both FCOS is and YOLO-v8 are anchor-free' should be 'Both FCOS and YOLO-v8 are anchor-free.'
- [Table 7] Several rows in Table 7 contain truncated numeric values with ellipses (e.g., '7204...' and '7086...'); these should be formatted as complete numbers or the table should state that truncated values are rounded for display.
Circularity Check
No significant circularity: the paper is a data/benchmark contribution whose statistics and baseline results are independent measurements, not derivations from their own outputs.
full rationale
The paper is a dataset and benchmark contribution, not a derivation chain in which an output quantity is defined in terms of the quantity it predicts. The central statistics (38,116 instances, 139 classes, 4,712 images, density and spatial-distribution properties) are direct measurements of the released annotations and of the images, so there is no self-definitional reduction. The baseline results in Tables 4-6 and the class-wise analysis in Table 7 are empirical evaluations of standard detectors on a fixed train/test split; they involve no fitted parameter that is afterwards renamed as a prediction, and the test AP numbers are not used to construct the dataset statistics. The supercategory evaluation in Table 6 is a different scoring protocol applied to the same detector outputs, not a quantity that was fitted to produce those outputs. Self-citations to the earlier ODOR challenge [10] and to SniffyArt [7] are contextual references to prior project versions; the central claim that ODOR is a large, dense, fine-grained artwork detection benchmark does not depend on those citations because the dataset itself is released and the baseline numbers are reproducible measurements. The acknowledged label ambiguity in Section 5 ('an inherent level of ambiguity persists') is a quality and validity limitation, but it is not a circularity: no claim reduces by construction to the annotation process, and the absence of inter-annotator agreement statistics is a correctness/evidence concern, not a circular-reasoning concern. The comparison to PeopleArt, IconArt, DEArt, PoPArt, Human-Art, and SniffyArt is an external empirical comparison based on reported dataset statistics, not a self-referential derivation. Therefore the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (3)
- minimum instance threshold =
16 instances across at least 3 images
- train/test split ratio =
90:10
- supercategory grouping rule =
pragmatic visual similarity, e.g., whales grouped as fish-like animals
assumptions (4)
- domain assumption The Odeuropa controlled vocabularies [59] operationalize 'olfactory references' in art, and the smell-related query terms used for collection define a valid target distribution for detection benchmarks.
- domain assumption A fixed two-level taxonomy can meaningfully categorize historical objects despite artistic abstraction, historical change, and borderline cases.
- domain assumption Expert review of crowd annotations yields accurate labels for 38,116 boxes across 139 fine-grained classes.
- standard math The COCO evaluation protocol (AP over IoU 0.50-0.95, COCO size bins) is the appropriate metric for quantifying benchmark difficulty.
Cite this review
Pith. "Pith review of Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset." pith.science (2026). https://pith.science/paper/3RNN5QZE
@misc{pith2026250708384,
author = {Pith},
title = {Pith review of: Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset},
year = {2026},
howpublished = {\url{https://pith.science/paper/3RNN5QZE}},
note = {Machine review of arXiv:2507.08384}
}
read the original abstract
Real-world applications of computer vision in the humanities require algorithms to be robust against artistic abstraction, peripheral objects, and subtle differences between fine-grained target classes. Existing datasets provide instance-level annotations on artworks but are generally biased towards the image centre and limited with regard to detailed object classes. The proposed ODOR dataset fills this gap, offering 38,116 object-level annotations across 4712 images, spanning an extensive set of 139 fine-grained categories. Conducting a statistical analysis, we showcase challenging dataset properties, such as a detailed set of categories, dense and overlapping objects, and spatial distribution over the whole image canvas. Furthermore, we provide an extensive baseline analysis for object detection models and highlight the challenging properties of the dataset through a set of secondary studies. Inspiring further research on artwork object detection and broader visual cultural heritage studies, the dataset challenges researchers to explore the intersection of object recognition and smell perception.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
H. Cai, Q. Wu, T. Corradi, P. Hall, The Cross-Depiction Problem: Computer Vision Algorithms for Recognising Objects in Artwork and in Photographs, arXiv preprint arXiv:1505.00110 (2015)
arXiv 2015
-
[2]
Madhu, A
P. Madhu, A. Villar-Corrales, R. Kosti, T. Bendschus, C. Reinhardt, P. Bell, A. Maier, V. Christlein, Enhancing Human Pose Estimation in Ancient Vase Paintings via Perceptually-grounded Style Transfer Learning, ACM Journal on Computing and Cultural Heritage 16 (1) (2022) 1–17
2022
-
[3]
Strezoski, M
G. Strezoski, M. Worring, OmniArt: A Large-scale Artistic Benchmark, ACM Transactions on Multimedia Computing, Communications, and Ap- plications (TOMM) 14 (4) (2018) 1–21
2018
-
[4]
Westlake, H
N. Westlake, H. Cai, P. Hall, Detecting People in Artwork with CNNs, in: Computer Vision–ECCV 2016 Workshops: Amsterdam, The Netherlands, October 8-10 and 15-16, 2016, Proceedings, Part I 14, Springer, 2016, pp. 825–841
2016
-
[5]
Gonthier, Y
N. Gonthier, Y. Gousseau, S. Ladjal, O. Bonfait, Weakly Supervised Object Detection in Artworks, in: L. Leal-Taix´ e, S. Roth (Eds.), Computer Vision – ECCV 2018 Workshops, Springer International Publishing, Cham, 2019, pp. 692–709
2018
-
[6]
Reshetnikov, M.-C
A. Reshetnikov, M.-C. Marinescu, J. M. Lopez, DEArt: Dataset of European Art, in: European Conference on Computer Vision, Springer, 2022, pp. 218– 233
2022
-
[7]
M. Zinnen, A. Hussian, H. Tran, P. Madhu, A. Maier, V. Christlein, Sniff- yArt: The Dataset of Smelling Persons, in: Proceedings of the 5th Workshop on AnalySis, Understanding and ProMotion of HeritAge Contents, SUMAC ’23, Association for Computing Machinery, New York, NY, USA, 2023, p. 49–58. doi:10.1145/3607542.3617357. URL https://doi.org/10.1145/36075...
-
[8]
X. Ju, A. Zeng, J. Wang, Q. Xu, L. Zhang, Human-Art: A versatile human- centric dataset bridging natural and artificial scenes, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 618–629
2023
Show all 92 references
-
[9]
Gupta, P
A. Gupta, P. Dollar, R. Girshick, L VIS: A Dataset for Large Vocabulary Instance Segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 5356–5364
2019
-
[10]
Zinnen, P
M. Zinnen, P. Madhu, R. Kosti, P. Bell, A. Maier, V. Christlein, ODOR: The ICPR2022 Odeuropa Challenge on Olfactory Object Recognition, in: 2022 26th International Conference on Pattern Recognition (ICPR), IEEE, 2022, pp. 4989–4994. 26
2022
-
[11]
S. Zhao, A. Akda˘ g Salah, A. A. Salah, Automatic Analysis of Human Body Representations in Western Art, in: Computer Vision–ECCV 2022 Workshops: Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part I, Springer, 2023, pp. 282–297
2022
-
[12]
Kadish, S
D. Kadish, S. Risi, A. S. Løvlie, Improving Object Detection in Art Images Using Only Style Transfer, in: 2021 International Joint Conference on Neural Networks (IJCNN), IEEE, 2021, pp. 1–8
2021
-
[13]
P. Hall, H. Cai, Q. Wu, T. Corradi, Cross-depiction problem: Recognition and synthesis of photographs and artwork, Computational Visual Media 1 (2015) 91–103
2015
-
[14]
Cetinic, J
E. Cetinic, J. She, Understanding and Creating Art with AI: Review and Outlook, ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) 18 (2) (2022) 1–22
2022
-
[15]
Y. Lu, C. Guo, X. Dai, F.-Y. Wang, Data-efficient image captioning of fine art paintings via virtual-real semantic alignment training, Neurocomputing 490 (2022) 163–180
2022
-
[16]
Sabatelli, M
M. Sabatelli, M. Kestemont, W. Daelemans, P. Geurts, Deep Transfer Learning for Art Classification Problems, in: L. Leal-Taix´ e, S. Roth (Eds.), Computer Vision – ECCV 2018 Workshops, Springer International Publish- ing, Cham, 2019, pp. 631–646
2018
-
[17]
Gonthier, Y
N. Gonthier, Y. Gousseau, S. Ladjal, An analysis of the transfer learning of convolutional neural networks for artistic images, in: Pattern Recognition. ICPR International Workshops and Challenges: Virtual Event, January 10–15, 2021, Proceedings, Part III, Springer, 2021, pp. 546–561
2021
-
[18]
Zinnen, P
M. Zinnen, P. Madhu, P. Bell, A. Maier, V. Christlein, Transfer Learn- ing for Olfactory Object Detection, in: Digital Humanities Conference, 2022, Alliance of Digital Humanities Organizations, 2022, pp. 409–413, https://arxiv.org/abs/2301.09906
2022 arXiv
-
[19]
W. Zhao, W. Jiang, X. Qiu, Big Transfer Learning for Fine Art Classification, Computational Intelligence and Neuroscience 2022 (2022)
2022
-
[20]
Cheng, J
G. Cheng, J. Han, P. Zhou, D. Xu, Learning Rotation-Invariant and Fisher Discriminative Convolutional Neural Networks for Object Detection, IEEE Transactions on Image Processing 28 (1) (2019) 265–278. doi:10.1109/TI P.2018.2867198
2019
-
[21]
Russakovsky, J
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al., Imagenet Large Scale Visual Recognition Challenge, International Journal of Computer Vision 115 (2015) 211–252. 27
2015
-
[22]
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll´ ar, C. L. Zitnick, Microsoft COCO: Common objects in context, in: Com- puter Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, Springer, 2014...
2014
-
[23]
Kuznetsova, H
A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov, et al., The Open Images Dataset v4: Unified image classification, object detection, and visual rela- tionship detection at scale, International Journal of ...
2020
-
[24]
Cheng, X
G. Cheng, X. Xie, J. Han, L. Guo, G.-S. Xia, Remote sensing image scene classification meets deep learning: Challenges, methods, benchmarks, and opportunities, IEEE Journal of Selected Topics in Applied Earth Observa- tions and Remote Sensing 13 (2020) 3735–3756
2020
-
[25]
Cheng, C
G. Cheng, C. Lang, M. Wu, X. Xie, X. Yao, J. Han, Feature Enhancement Network for Object Detection in Optical Remote Sensing Images, Journal of Remote Sensing (2021)
2021
-
[26]
Crowley, A
E. Crowley, A. Zisserman, The State of the Art: Object Retrieval in Paintings using Discriminative Regions, in: Proceedings of the British Machine Vision Conference. BMV A Press, 2014
2014
-
[27]
E. J. Crowley, A. Zisserman, In Search of Art, in: Computer Vision- ECCV 2014 Workshops: Zurich, Switzerland, September 6-7 and 12, 2014, Proceedings, Part I 13, Springer, 2015, pp. 54–70
2014
-
[28]
E. J. Crowley, A. Zisserman, The Art of Detection, in: Computer Vision– ECCV 2016 Workshops: Amsterdam, The Netherlands, October 8-10 and 15-16, 2016, Proceedings, Part I 14, Springer, 2016, pp. 721–737
2016
-
[29]
Girshick, Fast R-CNN, in: Proceedings of the IEEE international confer- ence on computer vision, 2015, pp
R. Girshick, Fast R-CNN, in: Proceedings of the IEEE international confer- ence on computer vision, 2015, pp. 1440–1448
2015
-
[30]
Madhu, A
P. Madhu, A. Meyer, M. Zinnen, L. M¨ uhrenberg, D. Suckow, T. Bendschus, C. Reinhardt, P. Bell, U. Verstegen, R. Kosti, et al., One-shot Object De- tection in Heterogeneous Artwork Datasets, in: 2022 Eleventh International Conference on Image Processing Theory, Tools and Appli...
2022
-
[31]
Madhu, T
P. Madhu, T. Marquart, R. Kosti, D. Suckow, P. Bell, A. Maier, V. Christlein, ICC++: Explainable feature learning for art history using image composi- tions, Pattern Recognition 136 (2023) 109153. doi:https://doi.org/10 .1016/j.patcog.2022.109153. URL https://www.sciencedirect...
2023
-
[32]
Madhu, T
P. Madhu, T. Marquart, R. Kosti, P. Bell, A. Maier, V. Christlein, Un- derstanding Compositional Structures in Art Historical Images Using Pose and Gaze Priors: Towards Scene Understanding in Digital Art History, in: Computer Vision–ECCV 2020 Workshops: Glasgow, UK, August 23–...
2020
-
[33]
P. Bell, L. Impett, The Choreography of the Annunciation through a Computational Eye, Histoire de l’art 34 (87) (2021) 01–06
2021
-
[34]
Bernasconi, GAB-Gestures for Artworks Browsing, in: 27th International Conference on Intelligent User Interfaces, 2022, pp
V. Bernasconi, GAB-Gestures for Artworks Browsing, in: 27th International Conference on Intelligent User Interfaces, 2022, pp. 50–53
2022
-
[35]
Impett, F
L. Impett, F. Moretti, Totentanz. Operationalizing Aby Warburg’s Pathos- formeln, Tech. rep., Stanford Literary Lab (2017)
2017
-
[36]
Impett, Analyzing Gesture in Digital Art History, in: The Routledge Companion to Digital Humanities and Art History, Routledge, 2020, pp
L. Impett, Analyzing Gesture in Digital Art History, in: The Routledge Companion to Digital Humanities and Art History, Routledge, 2020, pp. 386–407
2020
-
[37]
C. M. Becker, Aby Warburg’s Pathosformel as Methodological Paradigm, The Journal of Art Historiography 9 (2013) 9–CB1
2013
-
[38]
S. Kim, J. Park, J. Bang, H. Lee, Seeing is Smelling: Localizing Odor- Related Objects in Images, in: Proceedings of the 9th Augmented Human International Conference, 2018, pp. 1–9
2018
-
[39]
Y. Eda, H. Matsukura, Y. Nozaki, M. Sakamoto, Detection of odor-related objects in images based on everyday odors in Japan, in: Proceedings of the AAAI Spring Symposium: Socially Responsible AI for Well-being, 2023, pp. 59–60
2023
-
[40]
Rodr ´ ıguez-Ortega, Image Processing and Computer Vision in the Field of Art History, in: The Routledge Companion to Digital Humanities and Art History, Routledge, 2020, pp
N. Rodr ´ ıguez-Ortega, Image Processing and Computer Vision in the Field of Art History, in: The Routledge Companion to Digital Humanities and Art History, Routledge, 2020, pp. 338–357
2020
-
[41]
N¨ aslund Dahlgren, A
A. N¨ aslund Dahlgren, A. Wasielewski, Cultures of Digitization: A Historio- graphic Perspective on Digital Art History, Visual Resources 36 (4) (2020) 339–359
2020
-
[42]
S. Lang, B. Ommer, Reflecting on How Artworks Are Processed and Ana- lyzed by Computer Vision, in: Proceedings of the European Conference on Computer Vision (ECCV) Workshops, 2018
2018
-
[43]
Appadurai, The Social Life of Things, Tech
A. Appadurai, The Social Life of Things, Tech. rep., Cambridge University Press (1988)
1988
-
[44]
Hicks, M
D. Hicks, M. C. Beaudry, The Oxford handbook of material culture studies, OUP Oxford, 2010. 29
2010
-
[45]
M. J. Van Zuijlen, H. Lin, K. Bala, S. C. Pont, M. W. Wijntjes, Materials In Paintings (MIP): An interdisciplinary dataset for perception, art history, and computer vision, Plos one 16 (8) (2021) e0255109
2021
-
[46]
Leemans, W
I. Leemans, W. de Vries, Wind Trade: How the Concept of Wind Came to Embody Speculation in the Dutch Republic, The Journal of Modern History 94 (2) (2022) 288–325
2022
-
[47]
Everingham, A
M. Everingham, A. Zisserman, C. K. Williams, L. Van Gool, M. Allan, C. M. Bishop, O. Chapelle, N. Dalal, T. Deselaers, G. Dork´ o, et al., The 2005 PASCAL visual object classes challenge, in: Machine Learning Challenges. Evaluating Predictive Uncertainty, Visual Object Classif...
2005
-
[48]
Madhu, T
P. Madhu, T. Marquart, R. Kosti, D. Suckow, P. Bell, A. Maier, V. Christlein, Icc++: Explainable feature learning for art history using image composi- tions, Pattern Recognition 136 (2023) 109153
2023
-
[49]
Garcia, G
N. Garcia, G. Vogiatzis, How to Read Paintings: Semantic Art Understand- ing with Multi-Modal Retrieval, in: Proceedings of the European Conference on Computer Vision (ECCV) Workshops, 2018
2018
-
[50]
Garcia, C
N. Garcia, C. Ye, Z. Liu, Q. Hu, M. Otani, C. Chu, Y. Nakashima, T. Mi- tamura, A Dataset and Baselines for Visual Question Answering on Art, in: Computer Vision–ECCV 2020 Workshops: Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16, Springer, 2020, pp. 92–108
2020
-
[51]
M. J. Wilber, C. Fang, H. Jin, A. Hertzmann, J. Collomosse, S. Belongie, BAM! the Behance Artistic Media Dataset for Recognition Beyond Photog- raphy, in: Proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 1202–1211
2017
-
[52]
Wallace, D
A. Wallace, D. McCarthy, Survey of GLAM Open Access Policy and Practice, https://docs.google.com/document/d/15U__Z50WCUM_OWQ9HKLvLMlk cMoCN68FLVl9OKJQ8yY/edit?usp=sharing, accessed: 2023-02-02 (2018)
2018
-
[53]
Schneider, R
S. Schneider, R. Vollmer, Poses of People in Art: A Data Set for Human Pose Estimation in Digital Art History, arXiv preprint arXiv:2301.05124 (2023)
2023 arXiv
-
[54]
Gonthier, IconArt Dataset (Oct
N. Gonthier, IconArt Dataset (Oct. 2018). URL https://doi.org/10.5281/zenodo.4737435
2018 doi
-
[55]
Marinescu, A
M.-C. Marinescu, A. Reshetnikov, J. M. L´ opez, Improving Object Detection in Paintings based on Time Contexts, in: 2020 International Conference on Data Mining Workshops (ICDMW), IEEE, 2020, pp. 926–932. 30
2020
-
[56]
S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, J. Sun, Objects365: A Large-scale, High-quality Dataset for Object Detection, in: Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 8430–8439
2019
-
[57]
L. D. Couprie, Iconclass, a device for the iconographical analysis of art objects, Museum International 30 (3-4) (1978) 194–198
1978
-
[58]
Brandhorst, E
H. Brandhorst, E. Posthumus, Iconclass: a key to collaboration in the digital humanities, in: The Routledge Companion to Medieval Iconography, Routledge, 2016, pp. 201–218
2016
-
[59]
Lisena, D
P. Lisena, D. Schwabe, M. van Erp, R. Troncy, W. Tullett, I. Leemans, L. Marx, S. C. Ehrich, Capturing the Semantics of Smell: The Odeuropa Data Model for Olfactory Heritage Information, in: The Semantic Web: 19th International Conference, ESWC 2022, Hersonissos, Crete, Greece...
2022
-
[60]
G. A. Miller, WordNet: a Lexical Database for English, Communications of the ACM 38 (11) (1995) 39–41
1995
-
[61]
Redmon, A
J. Redmon, A. Farhadi, YOLO9000: better, faster, stronger, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 7263–7271
2017
-
[62]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sas- try, A. Askell, P. Mishkin, J. Clark, et al., Learning Transferable Visual Models from Natural Language Supervision, in: International Conference on Machine Learning, PMLR, 2021, pp. 8748–8763
2021
-
[63]
Kamath, M
A. Kamath, M. Singh, Y. LeCun, I. Misra, G. Synnaeve, N. Carion, MDETR – Modulated Detection for End-to-End Multi-Modal Understand- ing, arXiv:2104.12763 [cs]ArXiv: 2104.12763 (Apr. 2021). URL http://arxiv.org/abs/2104.12763
2021 arXiv
-
[64]
L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang, K.-W. Chang, J. Gao, Grounded Language-Image Pre-training, in: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, New Orleans, LA, USA, 2022, pp. 109...
2022
-
[65]
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu, L. Zhang, Grounding DINO: Marrying DINO with Grounded Pre- Training for Open-Set Object Detection, arXiv:2303.05499 [cs] (Mar. 2023). URL http://arxiv.org/abs/2303.05499
2023 arXiv
-
[66]
C. Xie, Z. Zhang, Y. Wu, F. Zhu, R. Zhao, S. Liang, Described Ob- ject Detection: Liberating Object Detection with Flexible Expressions, 31 arXiv:2307.12813 [cs] (Oct. 2023). URL http://arxiv.org/abs/2307.12813
2023 arXiv
-
[67]
Everingham, L
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, A. Zisserman, The pascal visual object classes (VOC) challenge, International journal of computer vision 88 (2009) 303–308
2009
-
[68]
Pont-Tuset, L
J. Pont-Tuset, L. Van Gool, Boosting object proposals: From PASCAL to COCO, in: Proceedings of the IEEE international conference on computer vision, 2015, pp. 1546–1554
2015
-
[69]
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, A. Torralba, Scene Parsing through ADE20K Dataset, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 633–641
2017
-
[70]
Y. Liu, P. Sun, N. Wergeles, Y. Shang, A survey and performance evaluation of deep learning methods for small object detection, Expert Systems with Applications 172 (2021) 114602
2021
-
[71]
van Erp, W
M. van Erp, W. Tullett, V. Christlein, T. Ehrhart, A. H¨ urriyeto˘ glu, I. Lee- mans, P. Lisena, S. Menini, D. Schwabe, S. Tonelli, et al., More than the Name of the Rose: How to Make Computers Read, See, and Organize Smells, The American Historical Review 128 (1) (2023) 335–369
2023
-
[72]
S. G. Magn´ usson, I. M. Szij´ art´ o, What is Microhistory?: Theory and Practice, Routledge, 2013
2013
-
[73]
H. Ali, T. Paccosi, S. Menini, Z. Mathias, L. Pasquale, A. Kiymet, T. Rapha¨ el, M. van Erp, MUSTI-Multimodal Understanding of Smells in Texts and Images at MediaEval 2022, in: Proceedings of MediaEval 2022 CEUR Workshop, 2022
2022
-
[74]
Howes, Sensual Relations: Engaging the Senses in Culture and Social Theory, University of Michiga/n Press, 2010
D. Howes, Sensual Relations: Engaging the Senses in Culture and Social Theory, University of Michiga/n Press, 2010
2010
-
[75]
Tullett, et al., Smell, History, and Heritage, The American Historical Review 127 (1) (2022) 261–309
W. Tullett, et al., Smell, History, and Heritage, The American Historical Review 127 (1) (2022) 261–309
2022
-
[76]
K. He, X. Zhang, S. Ren, J. Sun, Deep Residual Learning for Image Recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778
2016
-
[77]
S. Xie, R. Girshick, P. Doll´ ar, Z. Tu, K. He, Aggregated Residual Transfor- mations for Deep Neural Networks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 1492–1500
2017
-
[78]
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, B. Guo, Swin Transformer: Hierarchical Vision Transformer using Shifted Windows, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 10012–10022. 32
2021
-
[79]
Ridnik, E
T. Ridnik, E. Ben-Baruch, A. Noy, L. Zelnik-Manor, Imagenet-21k Pretrain- ing for the Masses, arXiv preprint arXiv:2104.10972 (2021)
2021 arXiv
-
[80]
J. Yang, C. Li, X. Dai, J. Gao, Focal Modulation Networks, Advances in Neural Information Processing Systems 35 (2022) 4203–4217
2022
-
[81]
S. Ren, K. He, R. Girshick, J. Sun, Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks, Advances in neural information processing systems 28 (2015)
2015
-
[82]
K. He, G. Gkioxari, P. Doll´ ar, R. Girshick, Mask R-CNN, in: Proceedings of the IEEE international conference on computer vision, 2017, pp. 2961–2969
2017
-
[83]
T.-Y. Lin, P. Doll´ ar, R. Girshick, K. He, B. Hariharan, S. Belongie, Feature Pyramid Networks for Object Detection, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2117– 2125
2017
-
[84]
Z. Cai, N. Vasconcelos, Cascade R-CNN: Delving into High Quality Object Detection, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 6154–6162
2018
-
[85]
Papers with Code, COCO object detection leaderboard, https://papers withcode.com/sota/object-detection-on-coco , accessed: February 19, 2023 (2021)
2021
-
[86]
Zhang, F
H. Zhang, F. Li, S. Liu, L. Zhang, H. Su, J. Zhu, L. Ni, H. Shum, DINO: DETR with Improved Denoising Anchor Boxes for End-to-End Object Detection, in: International Conference on Learning Representations, 2022
2022
-
[87]
W. Wang, J. Dai, Z. Chen, Z. Huang, Z. Li, X. Zhu, X. Hu, T. Lu, L. Lu, H. Li, et al., InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions, arXiv preprint arXiv:2211.05778 (2022)
2022 arXiv
-
[88]
Carion, F
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, S. Zagoruyko, End-to-End Object Detection with Transformers, in: Computer Vision– ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part I 16, Springer, 2020, pp. 213–229
2020
-
[89]
Z. Tian, C. Shen, H. Chen, T. He, FCOS: Fully Convolutional One-Stage Object Detection, in: 2019 IEEE/CVF International Conference on Com- puter Vision (ICCV), IEEE, Seoul, Korea (South), 2019, pp. 9626–9635. doi:10.1109/ICCV.2019.00972. URL https://ieeexplore.ieee.org/documen...
2019
-
[90]
X. Xie, C. Lang, S. Miao, G. Cheng, K. Li, J. Han, Mutual-Assistance Learning for Object Detection, IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (12) (2023) 15171–15184. doi:10.1109/TPAMI.20 23.3319634. URL https://ieeexplore.ieee.org/document/10265160/ 33
2023
-
[91]
Jocher, A
G. Jocher, A. Chaurasia, J. Qiu, YOLO by Ultralytics, https://github.c om/ultralytics/ultralytics (1 2023). 34 Appendix A. Image Credits Image Credits Figure 1a Detail from: Still Life with Bouquet of Flowers . Jan Brueghel the Elder. 1610 –
2023
-
[1625]
St¨ adel Museum, Frankfurt am Main
Oil on copper. St¨ adel Museum, Frankfurt am Main. https://www.staedelm useum.de/go/ds/540. Figure 1b Detail from: Village in the Snow . David Teniers (II). 1625 – 1690. Oil on panel. RKD – Netherlands Institute for Art History, RKDImages(290572). Figure 1c Detail from: Man Sm...
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.