Pith. sign in

REVIEW 4 major objections 8 minor 94 references

Open-World Panoptic Segmentation

T0 review · 4 major / 8 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A single fully convolutional network can, at test time, discover new semantic categories and segment individual instances of them, while keeping competitive accuracy on its training classes.

desk verdict PANIC is a genuinely useful benchmark and the task formulation helps the field, but the paper's headline claim about discovering novel categories is not yet backed by its own evaluation. read the letter →

arxiv 2412.12740 v1 pith:EPLQ2XEI submitted 2024-12-17 cs.CV cs.RO

classification cs.CVcs.RO
keywords open-worldpanopticsegmentationanomalynovelclassdiscoverydescriptorscontrastivelearninginstancebenchmarkdatasetautonomousdriving
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces the task of open-world panoptic segmentation: at test time, a vision system should not only flag objects it has never seen but also group them into new semantic categories and separate them into individual instances, all without retraining. The authors propose Con2MAV, a fully convolutional network with three decoders, and claim it is the first approach and the first benchmark with a public competition for this task. Con2MAV builds a fixed-dimensional descriptor for every class it was trained on, then uses those descriptors to open new classes when novel objects appear; a separate instance decoder locates each object within the anomalous area. The paper reports strong results across five datasets, covering driving, everyday, and underwater scenes, while keeping closed-world accuracy close to a model trained without any open-world component. It also releases PANIC, an 800-image driving benchmark with more than 50 unknown classes and over 4000 instances, with a hidden test set.

What carries the argument

The load-bearing object is the pre-logit class descriptor. Before the final convolution, every known class $k$ is represented by a running mean $\mu_k$ and variance $\sigma_k^2$ in a fixed-dimensional feature space $\mathbb{R}^D$, built from true-positive pixels only (Eqs. 2–3). The feature loss (Eq. 4) pulls each pixel's pre-logit toward its class's descriptor; the contrastive loss (Eq. 7) scatters class means on the unit sphere; and the objectosphere loss (Eq. 8) drives the norm of known-class features above 1 and unknown features toward 0. At test time, the squared exponential kernel (Eq. 18) scores a pixel's fit to each known Gaussian, and a $1\sigma$ bound decides "unknown"; the first unknown pixel seeds a new descriptor that is updated with a running mean and variance, so newly discovered classes stay consistent across images. The instance decoder's offset predictions are clustered with HDBSCAN only inside "thing" areas, and the semantic prediction filters clusters so pixels from different classes are not merged.

What would settle it

Train the same architecture on Cityscapes after removing all unlabeled background pixels from the objectosphere loss (or on a fully labeled dataset with no such background), evaluate on the PANIC hidden test set, and compare anomaly AUPR and unknown-class mIoU: if performance collapses, the background signal is load-bearing; if it holds, the method does not actually need unlabeled training pixels.

Watch

Extended reading notes

Core claim

The central claim is that discovering new semantic classes and new object instances at test time does not require out-of-distribution training data, generative models, or language models; a fully convolutional architecture with carefully chosen losses suffices. During training, the semantic decoder accumulates a running mean and variance of pre-logit features for each known class (Eqs. 2–3) and a feature loss pulls true-positive pixels toward those descriptors (Eq. 4). The contrastive decoder pushes known-class features onto the unit hypersphere while driving void or unlabeled feature norms toward zero (Eqs. 7–8), and the instance decoder predicts per-pixel offsets toward each instance centroid using a Lovász hinge plus divergence and curl regularizers (Eqs. 10–16). At test time a pixel is unknown only if it fails both a $1\sigma$ Gaussian-fit test against the known descriptors and a $1\sigma$ norm test from the contrastive decoder; the first unknown pixel's descriptor seeds a new class that evolves via a running average, and HDBSCAN clusters the offset predictions inside "thing" areas, with the semantic prediction filtering the resulting instances. The paper argues that this coupling enables consistent class discovery and instance segmentation together, and it reports top results on SegmentMeIfYouCan, BDDAnomaly, COCO, SUIM, and PANIC.

Load-bearing premise

The training signal for what counts as unknown comes from the unlabeled background areas of the Cityscapes images; if real-world anomalies look different from those background objects, the anomaly scores and new-class descriptors will be miscalibrated.

Editorial extensions

If this is right

  • A robot or vehicle can flag, categorize, and separate objects it has never seen without retraining or extra data, which is the main safety-relevant capability the paper targets.
  • The same training recipe transfers across domains: experiments on Cityscapes for driving, COCO for everyday scenes, and SUIM for underwater imagery all report competitive open-world and closed-world numbers.
  • Discovered classes persist: once an anomaly seeds a new descriptor, later appearances of the same category are matched to that descriptor, giving test-time consistency without incremental training.
  • The pre-logit design removes the fragile, manually tuned thresholds of the predecessor and also fixes the few-known-classes failure mode, shown by a roughly 19% mIoU gain on SUIM.
  • PANIC provides the first public, hidden-test-set benchmark for open-world panoptic segmentation in autonomous driving, with competitions for all four open-world tasks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The training signal for "unknown" is the unlabeled background pixels of Cityscapes; a deployment site whose novel objects look nothing like those background objects would likely see miscalibrated anomaly scores unless the detector is re-calibrated or augmented with out-of-distribution data.
  • Nothing in the method is specific to cameras: the losses and post-processing apply to any per-pixel feature backbone, so a LiDAR or RGB-D variant is a natural next step that the paper itself mentions as future work.
  • The open-world semantic evaluation matches each predicted class to the ground-truth class it overlaps most, which rewards purity but does not penalize splitting one true class into several predicted classes; adding a split-penalty metric would sharpen comparisons.
  • Because new classes are seeded from a single pixel's descriptor, an anomalous pixel that is unrepresentative of its category could create a noisy prototype; seeding classes from a small consensus cluster of mutually close anomalous descriptors is a testable robustness improvement.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 8 minor

Summary. The paper proposes Con2MAV, a fully convolutional network for open-world panoptic segmentation that extends the authors' previous ContMAV architecture. The method uses three decoders (semantic, contrastive, and instance) to segment known classes, detect unknown regions, discover novel semantic categories at test time, and segment instances within those categories. The paper also introduces PANIC, a new benchmark dataset for open-world segmentation in autonomous driving, with 800 images, 58 unknown classes, 4000+ instances, and public competitions for four tasks: anomaly segmentation, open-world semantic segmentation, open-set panoptic segmentation, and open-world panoptic segmentation. Experiments on SegmentMeIfYouCan, BDDAnomaly, COCO, SUIM, and PANIC report state-of-the-art results on several open-world tasks while retaining competitive closed-world performance. The paper also proposes a standardized nomenclature for open-world segmentation tasks.

Significance. If the results hold, the paper makes two valuable contributions. First, it extends open-world semantic segmentation to the panoptic setting, which is a natural and important step for autonomous perception. Second, the PANIC dataset and its public competitions address a real gap in benchmarks: existing anomaly datasets either lack semantic/instance annotations or have limited class diversity. The paper's approach is lightweight and the authors provide ablation studies showing the contribution of the pre-logit feature space and the instance-decoder losses. The explicit discussion of limitations and the attempt to unify task nomenclature are also useful. However, the evaluation has gaps that weaken the central claim of discovering truly novel categories, so the current version requires revision.

major comments (4)
  1. [5, Tables 8 and 12; Section 7] The PANIC validation set contains only 'known unknowns' (classes present in the Cityscapes void areas), while the hidden test set mixes these with 'unknown unknowns'; the paper reports only aggregate metrics on the hidden test set. Because the objectosphere loss (Eq. 8, Sec. 4.2) is trained on Cityscapes void areas, the learned anomaly detector is calibrated to the known-unknown distribution, and the aggregate mIoU/PQ figures could be dominated by the easy classes. The central claim that Con2MAV discovers truly novel categories at test time therefore needs separate results for the unknown-unknown subset (or per-class results) to be substantiated; the authors should report this breakdown or explicitly qualify the claim.
  2. [6.6, Table 12] For the newly introduced task of open-world panoptic segmentation, the paper reports only Con2MAV results and no comparative baseline. A simple baseline such as ContMAV (with an instance decoder or with the proposed clustering) or Mask2Anomaly would contextualize the 24.3% PQ and the claimed 'first approach' status. The absence of a baseline makes it impossible to assess whether the proposed modules are necessary for the task.
  3. [5.2.2] The proposed open-world IoU metric matches each predicted class to its best ground-truth class via argmax(row_i), which is permissive and rewards methods that over-segment. The paper does report homogeneity and completeness, but the headline mIoU numbers in Tables 6-9 and the state-of-the-art claims rely on this metric. The authors should add a stricter matching (e.g., Hungarian matching or a discussion of the effect of the argmax matching) and report the resulting numbers, or at least justify the choice with an analysis.
  4. [6, all tables] All reported results are single runs without error bars or significance tests. Given the stochasticity of training and the clustering post-processing, the claimed improvements (e.g., 'outperform ContMAV by 19% mIoU' on SUIM) may not be stable. At minimum, the main comparisons should include multiple seeds with standard deviations, or the authors should state that the differences are within run-to-run variability.
minor comments (8)
  1. [Abstract] The phrase 'object that have never been seen' should be 'objects that have never been seen'; also 'segmentaton' in the Introduction is a typo.
  2. [3.2] The definition of the open-world semantic mask Ms uses 0 for 'known' and 1..K for discovered categories, but the rest of the paper (e.g., Eq. (1), Table 2) uses 1..K for the known classes; this notation conflict should be resolved.
  3. [4.2, Eq. (11)] The notation e_p = p is confusing; it may be clearer to define the offset prediction o_p and write e_p = p + o_p.
  4. [Table 13] The last two columns of the ablation table are not labeled; the reader has to infer that they are mIoU and PQ.
  5. [5.2.2] Typo: 'evaluted' should be 'evaluated'.
  6. [7] Typo: 'intersting' should be 'interesting'.
  7. [Supplementary Fig. 9] Typo: 'extpected' should be 'expected'.
  8. [Table 3] The 'OoD' column header is not defined in the caption; define it.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the derivation chain is self-contained, and the cited prior work is a baseline rather than a load-bearing premise.

full rationale

The paper's central claims are supported by an end-to-end training and evaluation chain that does not reduce to its inputs. The semantic decoder trains with standard cross-entropy and a feature loss (Eqs. 1-5) whose class descriptors are running averages of ground-truth-filtered pre-logits; at test time, known/unknown decisions use a fixed 1-sigma bound and an objectosphere radius of 1 (Eqs. 8, 18-19), none of which are fit to the hidden test labels. The instance decoder uses standard offset, divergence, and curl losses (Eqs. 10-17), and the final clustering is post-processing evaluated on external and hidden splits. The heavy self-citations to ContMAV [75] provide the prior architecture and a baseline method, but Con2MAV's open-world panoptic contribution, including the new instance decoder and the PANIC benchmark, is evaluated against external datasets and a hidden test set. The stated limitation that the contrastive decoder relies on Cityscapes void areas (Sec. 7) is a distributional assumption, not a circular reduction: validation and test classes are not used to fit any parameter. No equation or fitted value is renamed as a prediction, so the circularity burden is not met.

Assumptions & free parameters 6 free parameters · 3 assumptions · 0 invented entities

The central claim rests on the availability of unlabeled/void areas in the training set to define the unknown distribution, and on the Gaussianity assumptions used for thresholding. These are not empirically verified, and the paper itself notes the void-area dependence as a limitation.

free parameters (6)
  • loss weights w1-w9 = 0.8, 0.2, 0.5, 0.5, 0.4, 0.2, 0.1, 0.2, 0.1
    Hand-chosen weights for the nine loss terms in Eqs. (5), (9), and (17). No sensitivity analysis is reported.
  • temperature tau = 0.1
    Temperature in the contrastive loss Eq. (7), taken from Chen et al. [11].
  • kernel width eta = not specified
    Width of the Gaussian in the instance offset soft mask Eq. (11). Affects clustering of instances.
  • 1sigma bound = 1
    Decision threshold for anomaly detection in post-processing, assuming Gaussian per-class feature distributions. Not fitted to test data.
  • objectosphere radius = 1
    Target norm for known features in Eq. (8). Argued to be required for synergy with contrastive loss.
  • HDBScan parameters = not specified
    Parameters for clustering offset predictions in the open-world panoptic post-processing. Not reported.
assumptions (3)
  • domain assumption Cityscapes void/unlabeled pixels are a valid training signal for unknown objects
    Section 4.2 uses void pixels as negatives in the objectosphere loss Eq. (8). Section 7 acknowledges this limitation.
  • ad hoc to paper Per-class pre-logit features are approximately Gaussian distributed
    The 1sigma bound for anomaly scoring assumes a multivariate normal with diagonal covariance. This is a modeling choice without validation.
  • ad hoc to paper Unknown-class features in the contrastive space are distributed as N(0,1)
    Section 4.3 states: 'we expect our unknown samples to be distributed as N(0,1)', used to set the 1sigma threshold.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Open-World Panoptic Segmentation." pith.science (2026). https://pith.science/paper/EPLQ2XEI

@misc{pith2026241212740,
  author       = {Pith},
  title        = {Pith review of: Open-World Panoptic Segmentation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EPLQ2XEI}},
  note         = {Machine review of arXiv:2412.12740}
}
read the original abstract

Perception is a key building block of autonomously acting vision systems such as autonomous vehicles. It is crucial that these systems are able to understand their surroundings in order to operate safely and robustly. Additionally, autonomous systems deployed in unconstrained real-world scenarios must be able of dealing with novel situations and object that have never been seen before. In this article, we tackle the problem of open-world panoptic segmentation, i.e., the task of discovering new semantic categories and new object instances at test time, while enforcing consistency among the categories that we incrementally discover. We propose Con2MAV, an approach for open-world panoptic segmentation that extends our previous work, ContMAV, which was developed for open-world semantic segmentation. Through extensive experiments across multiple datasets, we show that our model achieves state-of-the-art results on open-world segmentation tasks, while still performing competitively on the known categories. We will open-source our implementation upon acceptance. Additionally, we propose PANIC (Panoptic ANomalies In Context), a benchmark for evaluating open-world panoptic segmentation in autonomous driving scenarios. This dataset, recorded with a multi-modal sensor suite mounted on a car, provides high-quality, pixel-wise annotations of anomalous objects at both semantic and instance level. Our dataset contains 800 images, with more than 50 unknown classes, i.e., classes that do not appear in the training set, and 4000 object instances, making it an extremely challenging dataset for open-world segmentation tasks in the autonomous driving scenario. We provide competitions for multiple open-world tasks on a hidden test set. Our dataset and competitions are available at https://www.ipb.uni-bonn.de/data/panic.

Figures

Figures reproduced from arXiv: 2412.12740 by the authors.

Figure 1
Figure 1. Our proposed approach, Con2MAV, is able to tackle multiple open-world tasks and segment unknown objects and categories in multiple [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Our dataset, PANIC, provides pixel-wise annotations of unknown semantic categories and object instances of RGB images. The images [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. A schematic breakdown of the task discussed in this paper. [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Our network processes an RGB image via an encoder and three decoders and yields the final open-world panoptic segmentation result. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Qualitative results of our approach, Con2MAV, on open-world semantic segmentation on SUIM (top row) and PANIC (bottom row). [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: Qualitative results of our approach, Con2MAV, on open-set panoptic segmentation on COCO (top row), and PANIC (bottom row). The [PITH_FULL_IMAGE:figures/full_fig_p012_6.png]
Figure 7
Figure 7. Figure 7: Qualitative results of our approach, Con2MAV, on open-world panoptic segmentation on PANIC. The prediction mask is overlayed [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: A visual breakdown of the four open-world segmentation tasks discussed in this paper. Anomaly segmentation segment all anomalous [PITH_FULL_IMAGE:figures/full_fig_p017_8.png]
Figure 9
Figure 9. Figure 9: A visual breakdown of the interaction between contrastive and [PITH_FULL_IMAGE:figures/full_fig_p017_9.png]
Figure 10
Figure 10. Figure 10: Qualitative results on anomaly segmentation. We show the input RGB image on the left, the prediction of our previous approach [PITH_FULL_IMAGE:figures/full_fig_p019_10.png]
Figure 11
Figure 11. Figure 11: Qualitative results on open-world semantic segmentation. We show the input RGB image on the left, the prediction of our previous [PITH_FULL_IMAGE:figures/full_fig_p020_11.png]
Figure 12
Figure 12. Figure 12: Qualitative results on open-set panoptic segmentation. We show the input RGB image on the left, the anomaly prediction of Con2MAV [PITH_FULL_IMAGE:figures/full_fig_p021_12.png]
Figure 13
Figure 13. Figure 13: Qualitative results on open-world panoptic segmentation. We show the input RGB image on the left, the open-world semantic prediction [PITH_FULL_IMAGE:figures/full_fig_p022_13.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

94 extracted references · 77 canonical work pages

  1. [1]

    Maskomaly: Zero-shot mask anomaly segmentation,

    J. Ackermann, C. Sakaridis, and F. Yu, “Maskomaly: Zero-shot mask anomaly segmentation,” in Proc. of British Machine Vision Conference (BMVC), 2023. 10

  2. [2]

    A General and Adaptive Robust Loss Function,

    J. T. Barron, “A General and Adaptive Robust Loss Function,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2019. 7

  3. [3]

    SemanticKITTI: A Dataset for Semantic Scene Understand- ing of LiDAR Sequences,

    J. Behley, M. Garbade, A. Milioto, J. Quenzel, S. Behnke, C. Stachniss, and J. Gall, “SemanticKITTI: A Dataset for Semantic Scene Understand- ing of LiDAR Sequences,” in Proc. of the IEEE/CVF Intl. Conf. on Computer Vision (ICCV), 2019. 13

  4. [4]

    Towards open set deep networks,

    A. Bendale and T. E. Boult, “Towards open set deep networks,” in Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2016. 2, 8, 13

  5. [5]

    MVTec AD–A comprehensive real-world dataset for unsupervised anomaly detection,

    P. Bergmann, M. Fauser, D. Sattlegger, and C. Steger, “MVTec AD–A comprehensive real-world dataset for unsupervised anomaly detection,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2019. 3

  6. [6]

    Fishyscapes: A benchmark for safe semantic segmentation in autonomous driving,

    H. Blum, P. E. Sarlin, J. Nieto, R. Siegwart, and C. Cadena, “Fishyscapes: A benchmark for safe semantic segmentation in autonomous driving,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition Workshops, 2019. 3, 8

  7. [7]

    The Fishyscapes Benchmark: Measuring blind spots in semantic segmenta- tion,

    H. Blum, P. E. Sarlin, J. Nieto, R. Siegwart, and C. Cadena, “The Fishyscapes Benchmark: Measuring blind spots in semantic segmenta- tion,” Intl. Journal of Computer Vision (IJCV), vol. 129, pp. 3119–3135,

  8. [8]

    Inverseform: A loss function for structured boundary-aware segmentation,

    S. Borse, Y . Wang, Y . Zhang, and F. Porikli, “Inverseform: A loss function for structured boundary-aware segmentation,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) ,

Show all 94 references
  1. [9]

    Weightless neural networks for open set recognition,

    D. O. Cardoso, J. Gama, and F. M. Franc ¸a, “Weightless neural networks for open set recognition,” Machine Learning , vol. 106, no. 9-10, pp. 1547–1567, 2017. 2

  2. [10]

    SegmentMeIfYouCan: A Benchmark for Anomaly Segmentation,

    R. Chan, K. Lis, S. Uhlemeyer, H. Blum, S. Honari, R. Siegwart, P. Fua, M. Salzmann, and M. Rottmann, “SegmentMeIfYouCan: A Benchmark for Anomaly Segmentation,” in Proc. of the Conf. on Neural Information Processing Systems (NeurIPS), 2021. 1, 2, 3, 4, 7, 8, 9, 13, 19

  3. [11]

    A Simple Framework for Contrastive Learning of Visual Representations,

    T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A Simple Framework for Contrastive Learning of Visual Representations,” in Proc. of the Intl. Conf. on Machine Learning (ICML), 2020. 6, 10, 17

  4. [12]

    Panoptic-DeepLab: A Simple, Strong, and Fast Baseline for Bottom-Up Panoptic Segmentation,

    B. Cheng, M. D. Collins, Y . Zhu, T. Liu, T. S. Huang, H. Adam, and L.- C. Chen, “Panoptic-DeepLab: A Simple, Strong, and Fast Baseline for Bottom-Up Panoptic Segmentation,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2020. 1

  5. [13]

    Masked- attention mask transformer for universal image segmentation,

    B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar, “Masked- attention mask transformer for universal image segmentation,” inProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2022. 1, 13

  6. [14]

    Xception: Deep learning with depthwise separable convolu- tions,

    F. Chollet, “Xception: Deep learning with depthwise separable convolu- tions,” in cvpr, 2017. 4

  7. [15]

    Weakly and semi-supervised detection, segmentation and tracking of table grapes with limited and noisy data,

    T. A. Ciarfuglia, I. M. Motoi, L. Saraceni, M. Fawakherji, A. Sanfeliu, and D. Nardi, “Weakly and semi-supervised detection, segmentation and tracking of table grapes with limited and noisy data,” Computers and Electronics in Agriculture, vol. 205, p. 107624, 2023. 2

  8. [16]

    The Cityscapes dataset for semantic urban scene understanding,

    M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The Cityscapes dataset for semantic urban scene understanding,” in Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2016. 1, 3, 4, 7, 8

  9. [17]

    Outlier detection by en- sembling uncertainty with negative objectness,

    A. Deli ´c, M. Grci ´c, and S. Segvi ´c, “Outlier detection by en- sembling uncertainty with negative objectness,” arXiv preprint , vol. arXiv:2402.15374, 2024. 1, 10, 13

  10. [18]

    Learning confidence for out- of-distribution detection in neural networks,

    T. DeVries and G. W. Taylor, “Learning confidence for out- of-distribution detection in neural networks,” arXiv preprint , vol. arXiv:1802.04865, 2018. 10

  11. [19]

    Reducing network agnosto- phobia,

    A. R. Dhamija, M. G ¨unther, and T. Boult, “Reducing network agnosto- phobia,” in Proc. of the Conf. on Neural Information Processing Systems (NeurIPS), 2018. 2, 6, 17 15

  12. [20]

    The pascal visual object classes (voc) challenge,

    M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes (voc) challenge,” Intl. Journal of Computer Vision (IJCV), vol. 88, pp. 303–338, 2010. 9

  13. [21]

    Dropout as a Bayesian Approximation: Representing model uncertainty in deep learning,

    Y . Gal and Z. Ghahramani, “Dropout as a Bayesian Approximation: Representing model uncertainty in deep learning,” in Proc. of the Intl. Conf. on Machine Learning (ICML), 2016. 2, 10

  14. [22]

    Segmenting known objects and unseen unknowns without prior knowledge,

    S. Gasperini, A. Marcos-Ramiro, M. Schmidt, N. Navab, B. Busam, and F. Tombari, “Segmenting known objects and unseen unknowns without prior knowledge,” in Proc. of the IEEE/CVF Intl. Conf. on Computer Vision (ICCV), 2023. 2

  15. [23]

    Scaling open-vocabulary image segmentation with image-level labels,

    G. Ghiasi, X. Gu, Y . Cui, and T.-Y . Lin, “Scaling open-vocabulary image segmentation with image-level labels,” in Proc. of the Europ. Conf. on Computer Vision (ECCV), 2022. 3

  16. [24]

    Fast R-CNN,

    R. Girshick, “Fast R-CNN,” in Proc. of the IEEE Intl. Conf. on Computer Vision (ICCV), 2015. 1

  17. [25]

    Densehybrid: Hybrid anomaly detection for dense open-set recognition,

    M. Grci ´c, P. Bevandi ´c, and S. Segvi ´c, “Densehybrid: Hybrid anomaly detection for dense open-set recognition,” in Proc. of the Europ. Conf. on Computer Vision (ECCV), 2022. 10, 13

  18. [26]

    On advantages of mask-level recog- nition for outlier-aware segmentation,

    M. Grci ´c, J. Sari ´c, and S. Segvi ´c, “On advantages of mask-level recog- nition for outlier-aware segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023. 1

  19. [27]

    Disadvantages of using the area under the receiver operating characteristic curve to assess imaging tests: a discussion and proposal for an alternative approach,

    S. Halligan, D. G. Altman, and S. Mallett, “Disadvantages of using the area under the receiver operating characteristic curve to assess imaging tests: a discussion and proposal for an alternative approach,” European Radiology, vol. 25, pp. 932–939, 2015. 8

  20. [28]

    Mask r-cnn,

    K. He, G. Gkioxari, P. Doll ´ar, and R. Girshick, “Mask r-cnn,” in Proc. of the IEEE/CVF Intl. Conf. on Computer Vision (ICCV), 2017. 1

  21. [29]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2016. 4

  22. [30]

    Scaling out-of-distribution detection for real- world settings,

    D. Hendrycks, S. Basart, M. Mazeika, A. Zou, J. Kwon, M. Mostajabi, J. Steinhardt, and D. Song, “Scaling out-of-distribution detection for real- world settings,” Proc. of the Intl. Conf. on Machine Learning (ICML) ,

  23. [31]

    A baseline for detecting misclassified and out-of-distribution examples in neural networks,

    D. Hendrycks and K. Gimpel, “A baseline for detecting misclassified and out-of-distribution examples in neural networks,” in Proc. of the Intl. Conf. on Learning Representations (ICLR), 2017. 2, 10, 18

  24. [32]

    Bidirectional projec- tion network for cross dimension scene understanding,

    W. Hu, H. Zhao, L. Jiang, J. Jia, and T.-T. Wong, “Bidirectional projec- tion network for cross dimension scene understanding,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) ,

  25. [33]

    Beyond auroc & co. for evaluating out-of-distribution detection performance,

    G. Humblot-Renaux, S. Escalera, and T. B. Moeslund, “Beyond auroc & co. for evaluating out-of-distribution detection performance,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2023. 8

  26. [34]

    Exemplar-based open-set panoptic segmentation network,

    J. Hwang, S. W. Oh, J.-Y . Lee, and B. Han, “Exemplar-based open-set panoptic segmentation network,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2021. 2, 4, 13

  27. [35]

    Semantic segmentation of underwater imagery: Dataset and benchmark,

    M. J. Islam, C. Edge, Y . Xiao, P. Luo, M. Mehtaz, C. Morse, S. S. Enan, and J. Sattar, “Semantic segmentation of underwater imagery: Dataset and benchmark,” in Proc. of the IEEE/RSJ Intl. Conf. on Intelligent Robots and Systems (IROS), 2020. 1, 2, 9, 11, 20

  28. [36]

    Scaling up visual and vision-language representation learning with noisy text supervision,

    C. Jia, Y . Yang, Y . Xia, Y .-T. Chen, Z. Parekh, H. Pham, Q. Le, Y .- H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in Proc. of the Intl. Conf. on Machine Learning (ICML), 2021. 3

  29. [37]

    Adam: A Method for Stochastic Optimization,

    D. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,” in Proc. of the Intl. Conf. on Learning Representations (ICLR), 2015. 10

  30. [38]

    Panoptic feature pyramid networks,

    A. Kirillov, R. Girshick, K. He, and P. Doll ´ar, “Panoptic feature pyramid networks,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2019. 1

  31. [39]

    Panoptic segmentation,

    A. Kirillov, K. He, R. Girshick, C. Rother, and P. Doll ´ar, “Panoptic segmentation,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2019. 9

  32. [40]

    Opengan: Open-set recognition via open data generation,

    S. Kong and D. Ramanan, “Opengan: Open-set recognition via open data generation,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2021. 3

  33. [41]

    Virtual multi-view fusion for 3d semantic segmentation,

    A. Kundu, X. Yin, A. Fathi, D. Ross, B. Brewington, T. Funkhouser, and C. Pantofaru, “Virtual multi-view fusion for 3d semantic segmentation,” in Proc. of the Europ. Conf. on Computer Vision (ECCV), 2020. 2

  34. [42]

    Simple and scalable predictive uncertainty estimation using deep ensembles,

    B. Lakshminarayanan, A. Pritzel, and C. Blundell, “Simple and scalable predictive uncertainty estimation using deep ensembles,” Proc. of the Conf. on Neural Information Processing Systems (NeurIPS), 2017. 2

  35. [43]

    Open-vocabulary semantic segmentation with mask- adapted clip,

    F. Liang, B. Wu, X. Dai, K. Li, Y . Zhao, H. Zhang, P. Zhang, P. Vajda, and D. Marculescu, “Open-vocabulary semantic segmentation with mask- adapted clip,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2023. 3

  36. [44]

    Microsoft COCO: Common objects in context,

    T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll ´ar, and C. L. Zitnick, “Microsoft COCO: Common objects in context,” in Proc. of the Europ. Conf. on Computer Vision (ECCV), 2014. 1, 2, 3, 9, 18, 20, 21

  37. [45]

    Path aggregation network for instance segmentation,

    S. Liu, L. Qi, H. Qin, J. Shi, and J. Jia, “Path aggregation network for instance segmentation,” Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2018. 2

  38. [46]

    Energy-based out-of-distribution detection,

    W. Liu, X. Wang, J. Owens, and Y . Li, “Energy-based out-of-distribution detection,” in Proc. of the Conf. on Neural Information Processing Systems (NeurIPS), 2020. 2

  39. [47]

    Opening up open world tracking,

    Y . Liu, I. E. Zulfikar, J. Luiten, A. Dave, D. Ramanan, B. Leibe, A. Osep, and L. Leal-Taix ´e, “Opening up open world tracking,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) ,

  40. [48]

    Fully convolutional networks for semantic segmentation,

    J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2015. 1

  41. [49]

    Pixel-wise gradient uncertainty for convo- lutional neural networks applied to out-of-distribution segmentation,

    K. Maag and T. Riedlinger, “Pixel-wise gradient uncertainty for convo- lutional neural networks applied to out-of-distribution segmentation,” in Intl. Conf. on Computer Vision Theory and Applications (VISAPP), 2024. 2

  42. [50]

    High precision leaf instance segmentation for phenotyping in point clouds obtained under real field conditions,

    E. Marks, M. Sodano, F. Magistri, L. Wiesmann, D. Desai, R. Marcuzzi, J. Behley, and C. Stachniss, “High precision leaf instance segmentation for phenotyping in point clouds obtained under real field conditions,” IEEE Robotics and Automation Letters (RA-L) , vol. 8, pp. 4791–4798,

  43. [51]

    HDBScan: Hierarchical density based clustering

    L. McInnes, J. Healy, S. Astels et al., “HDBScan: Hierarchical density based clustering.” Journal of Open Source Software, vol. 2, p. 205, 2017. 7

  44. [52]

    Fast Instance and Semantic Segmentation Exploiting Local Connectivity, Metric Learning, and One- Shot Detection for Robotics,

    A. Milioto, L. Mandtler, and C. Stachniss, “Fast Instance and Semantic Segmentation Exploiting Local Connectivity, Metric Learning, and One- Shot Detection for Robotics,” inProc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA), 2019. 2

  45. [53]

    Bonnet: An Open-Source Training and Deployment Framework for Semantic Segmentation in Robotics using CNNs,

    A. Milioto and C. Stachniss, “Bonnet: An Open-Source Training and Deployment Framework for Semantic Segmentation in Robotics using CNNs,” in Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA), 2019. 1, 2

  46. [54]

    Real-time Semantic Segmen- tation of Crop and Weed for Precision Agriculture Robots Leveraging Background Knowledge in CNNs,

    A. Milioto, P. Lottes, and C. Stachniss, “Real-time Semantic Segmen- tation of Crop and Weed for Precision Agriculture Robots Leveraging Background Knowledge in CNNs,” in Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA), 2018. 2

  47. [55]

    Confidence prediction for lexicon-free ocr,

    N. Mor and L. Wolf, “Confidence prediction for lexicon-free ocr,” in Proc. of the IEEE Winter Conf. on Applications of Computer Vision (WACV), 2018. 2

  48. [56]

    Rba: Segmenting unknown regions rejected by all,

    N. Nayal, M. Yavuz, J. F. Henriques, and F. G ¨uney, “Rba: Segmenting unknown regions rejected by all,” inProc. of the IEEE/CVF Intl. Conf. on Computer Vision (ICCV), 2023. 1, 2, 10, 13

  49. [57]

    Ugains: Uncer- tainty guided anomaly instance segmentation,

    A. Nekrasov, A. Hermans, L. Kuhnert, and B. Leibe, “Ugains: Uncer- tainty guided anomaly instance segmentation,” in Proc. of the German Conf. on Pattern Recognition (GCPR), 2023. 2

  50. [58]

    Oodis: Anomaly instance segmentation benchmark,

    A. Nekrasov, R. Zhou, M. Ackermann, A. Hermans, B. Leibe, and M. Rottmann, “Oodis: Anomaly instance segmentation benchmark,” arXiv preprint, vol. arXiv:2406.11835, 2024. 3, 8

  51. [59]

    Instance Segmentation by Jointly Optimizing Spatial Embeddings and Clustering Bandwidth,

    D. Neven, B. D. Brabandere, M. Proesmans, and L. V . Gool, “Instance Segmentation by Jointly Optimizing Spatial Embeddings and Clustering Bandwidth,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2019. 6

  52. [60]

    Deep neural networks are easily fooled: High confidence predictions for unrecognizable images,

    A. Nguyen, J. Yosinski, and J. Clune, “Deep neural networks are easily fooled: High confidence predictions for unrecognizable images,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2015. 2

  53. [61]

    In defense of pre- trained imagenet architectures for real-time semantic segmentation of road-driving images,

    M. Orsic, I. Kreso, P. Bevandic, and S. Segvic, “In defense of pre- trained imagenet architectures for real-time semantic segmentation of road-driving images,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2019. 4

  54. [62]

    Lost and found: detecting small road hazards for self-driving vehicles,

    P. Pinggera, S. Ramos, S. Gehrig, U. Franke, C. Rother, and R. Mester, “Lost and found: detecting small road hazards for self-driving vehicles,” in Proc. of the IEEE/RSJ Intl. Conf. on Intelligent Robots and Systems (IROS), 2016. 3

  55. [63]

    Freeseg: Unified, universal and open-vocabulary image segmentation,

    J. Qin, J. Wu, P. Yan, M. Li, R. Yuxi, X. Xiao, Y . Wang, R. Wang, S. Wen, X. Pan et al. , “Freeseg: Unified, universal and open-vocabulary image segmentation,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2023. 3

  56. [64]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in Proc. of the Intl. Conf. on Machine Learning (ICML) , 2021. 3 16

  57. [65]

    Mask2anomaly: Mask transformer for universal open-set segmentation,

    S. N. Rai, F. Cermelli, B. Caputo, and C. Masone, “Mask2anomaly: Mask transformer for universal open-set segmentation,”IEEE Trans. on Pattern Analysis and Machine Intelligence (TPAMI), 2024. 1, 2, 4, 9, 13, 18

  58. [66]

    Faster r-cnn: Towards real- time object detection with region proposal networks,

    S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real- time object detection with region proposal networks,” IEEE Trans. on Pattern Analysis and Machine Intelligence (TPAMI) , vol. 39, no. 6, pp. 1137–1149, 2016. 1

  59. [67]

    Hierarchical approach for joint semantic, plant instance, and leaf instance segmentation in the agricultural domain,

    G. Roggiolani, M. Sodano, T. Guadagnino, F. Magistri, J. Behley, and C. Stachniss, “Hierarchical approach for joint semantic, plant instance, and leaf instance segmentation in the agricultural domain,” in Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA) , 2023. 1, 2

  60. [68]

    Erfnet: Ef- ficient residual factorized convnet for real-time semantic segmentation,

    E. Romera, J. M. Alvarez, L. M. Bergasa, and R. Arroyo, “Erfnet: Ef- ficient residual factorized convnet for real-time semantic segmentation,” IEEE Trans. on Intelligent Transportation Systems (ITS) , vol. 19, no. 1, pp. 263–272, 2018. 4

  61. [69]

    U-net: Convolutional networks for biomedical image segmentation,

    O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Proc. of the Intl. Conf. on Medical image computing and computer-assisted intervention (MICCAI),

  62. [70]

    V-measure: A Conditional Entropy- Based External Cluster Evaluation Measure,

    A. Rosenberg and J. Hirschberg, “V-measure: A Conditional Entropy- Based External Cluster Evaluation Measure,” in Proc. of the Joint Conf. on Empirical Methods in Natural Language Processing and Com- putational Natural Language Learning (EMNLP-CoNLL), 2007. 9

  63. [71]

    Multiresolution Knowledge Distillation for Anomaly Detection,

    M. Salehi, N. Sadjadi, S. Baselizadeh, M. H. Rohban, and H. R. Rabiee, “Multiresolution Knowledge Distillation for Anomaly Detection,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recog- nition (CVPR), 2021. 3

  64. [72]

    Bayesian Nonparametric Submodular Video Partition for Robust Anomaly Detection,

    H. Sapkota and Q. Yu, “Bayesian Nonparametric Submodular Video Partition for Robust Anomaly Detection,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2022. 2

  65. [73]

    Super-convergence: Very fast training of neural networks using large learning rates,

    L. N. Smith and N. Topin, “Super-convergence: Very fast training of neural networks using large learning rates,” in Artificial Intelligence and Machine Learning for Multi-Domain Operations Applications , vol. 11006, 2019, pp. 369–386. 10

  66. [74]

    Robust double-encoder network for rgb-d panoptic segmentation,

    M. Sodano, F. Magistri, T. Guadagnino, J. Behley, and C. Stachniss, “Robust double-encoder network for rgb-d panoptic segmentation,” in Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA) , 2023. 1, 2

  67. [75]

    Open- world semantic segmentation including class similarity,

    M. Sodano, F. Magistri, L. Nunes, J. Behley, and C. Stachniss, “Open- world semantic segmentation including class similarity,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) ,

  68. [76]

    Look closer to segment better: Boundary patch refinement for instance segmentation,

    C. Tang, H. Chen, X. Li, J. Li, Z. Zhang, and X. Hu, “Look closer to segment better: Boundary patch refinement for instance segmentation,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2021. 1

  69. [77]

    A deeper look at dataset bias,

    T. Tommasi, N. Patricia, B. Caputo, and T. Tuytelaars, “A deeper look at dataset bias,” Domain Adaptation in Computer Vision Applications , pp. 37–55, 2017. 2

  70. [78]

    Open-set recognition: A good closed-set classifier is all you need,

    S. Vaze, K. Han, A. Vedaldi, and A. Zisserman, “Open-set recognition: A good closed-set classifier is all you need,” in Proc. of the Intl. Conf. on Learning Representations (ICLR), 2021. 2

  71. [79]

    Toward reproducible version-controlled perception plat- forms: Embracing simplicity in autonomous vehicle dataset acquisition,

    I. Vizzo, B. Mersch, L. Nunes, L. Wiesmann, T. Guadagnino, and C. Stachniss, “Toward reproducible version-controlled perception plat- forms: Embracing simplicity in autonomous vehicle dataset acquisition,” in Proc. of the IEEE Intl. Conf. on Intelligent Transportation Systems ...

  72. [80]

    Out-of-distribution detection using an ensemble of self supervised leave- out classifiers,

    A. Vyas, N. Jammalamadaka, X. Zhu, D. Das, B. Kaul, and T. L. Willke, “Out-of-distribution detection using an ensemble of self supervised leave- out classifiers,” in Proc. of the Europ. Conf. on Computer Vision (ECCV),

  73. [81]

    Internimage: Exploring large-scale vision foundation models with deformable convolutions,

    W. Wang, J. Dai, Z. Chen, Z. Huang, Z. Li, X. Zhu, X. Hu, T. Lu, L. Lu, H. Li, X. Wang, and Y . Qiao, “Internimage: Exploring large-scale vision foundation models with deformable convolutions,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) ,

  74. [82]

    Openauc: Towards auc-oriented open-set recognition,

    Z. Wang, Q. Xu, Z. Yang, Y . He, X. Cao, and Q. Huang, “Openauc: Towards auc-oriented open-set recognition,”Proc. of the Conf. on Neural Information Processing Systems (NeurIPS), 2022. 8

  75. [83]

    Panoptic Segmentation with Partial Annotations for Agricultural Robots,

    J. Weyler, T. L ¨abe, J. Behley, and C. Stachniss, “Panoptic Segmentation with Partial Annotations for Agricultural Robots,” IEEE Robotics and Automation Letters (RA-L), vol. 9, no. 2, pp. 1660–1667, 2024. 5, 6, 7

  76. [84]

    PhenoBench–A Large Dataset and Benchmarks for Semantic Image Interpretation in the Agricultural Domain,

    J. Weyler, F. Magistri, E. Marks, Y . L. Chong, M. Sodano, G. Roggiolani, N. Chebrolu, C. Stachniss, and J. Behley, “PhenoBench–A Large Dataset and Benchmarks for Semantic Image Interpretation in the Agricultural Domain,” IEEE Trans. on Pattern Analysis and Machine Intelligenc...

  77. [85]

    Identifying unknown instances for autonomous driving,

    K. Wong, S. Wang, M. Ren, M. Liang, and R. Urtasun, “Identifying unknown instances for autonomous driving,” in Proc. of the Conf. on Robot Learning (CoRL), 2020. 2

  78. [86]

    Masqclip for open-vocabulary universal image segmentation,

    X. Xu, T. Xiong, Z. Ding, and Z. Tu, “Masqclip for open-vocabulary universal image segmentation,” in Proc. of the IEEE/CVF Intl. Conf. on Computer Vision (ICCV), 2023. 3

  79. [87]

    Codabench: Flexible, easy-to-use, and reproducible meta- benchmark platform,

    Z. Xu, S. Escalera, A. Pav ˜ao, M. Richard, W.-W. Tu, Q. Yao, H. Zhao, and I. Guyon, “Codabench: Flexible, easy-to-use, and reproducible meta- benchmark platform,” Patterns, vol. 3, no. 7, p. 100543, 2022. 8, 18

  80. [88]

    BDD100K: A diverse driving dataset for heterogeneous multitask learning,

    F. Yu, H. Chen, X. Wang, W. Xian, Y . Chen, F. Liu, V . Madhavan, and T. Darrell, “BDD100K: A diverse driving dataset for heterogeneous multitask learning,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2020. 3, 8

  81. [89]

    Wilddash-creating hazard-aware benchmarks,

    O. Zendel, K. Honauer, M. Murschitz, D. Steininger, and G. F. Dominguez, “Wilddash-creating hazard-aware benchmarks,” in Proc. of the Europ. Conf. on Computer Vision (ECCV), 2018. 3

  82. [90]

    A simple framework for open-vocabulary segmentation and detection,

    H. Zhang, F. Li, X. Zou, S. Liu, C. Li, J. Yang, and L. Zhang, “A simple framework for open-vocabulary segmentation and detection,” in Proc. of the IEEE/CVF Intl. Conf. on Computer Vision (ICCV), 2023. 3

  83. [91]

    Pyramid Scene Parsing Network,

    H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid Scene Parsing Network,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2017. 4

  84. [92]

    OmniAL: A Unified CNN Framework for Unsupervised Anomaly Localization,

    Y . Zhao, “OmniAL: A Unified CNN Framework for Unsupervised Anomaly Localization,” in Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2023. 3 Matteo Sodano is a PhD student at the Pho- togrammetry & Robotics Lab at the University of Bonn since Ja...

  85. [2022]

    1, 2, 3, 8, 9, 10, 19

  86. [2024]

    1, 2, 4, 5, 10, 11, 12, 17, 18, 19, 20

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.