Pith. sign in

REVIEW 6 cited by

Semantic Instance Segmentation with a Discriminative Loss Function

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1708.02551 v1 pith:7NXU7NYC submitted 2017-08-08 cs.CV cs.RO

classification cs.CVcs.RO
keywords functioninstancelosssegmentationnetworksimplediscriminativeencourages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Semantic instance segmentation remains a challenging task. In this work we propose to tackle the problem with a discriminative loss function, operating at the pixel level, that encourages a convolutional network to produce a representation of the image that can easily be clustered into instances with a simple post-processing step. The loss function encourages the network to map each pixel to a point in feature space so that pixels belonging to the same instance lie close together while different instances are separated by a wide margin. Our approach of combining an off-the-shelf network with a principled loss function inspired by a metric learning objective is conceptually simple and distinct from recent efforts in instance segmentation. In contrast to previous works, our method does not rely on object proposals or recurrent mechanisms. A key contribution of our work is to demonstrate that such a simple setup without bells and whistles is effective and can perform on par with more complex methods. Moreover, we show that it does not suffer from some of the limitations of the popular detect-and-segment approaches. We achieve competitive performance on the Cityscapes and CVPPP leaf segmentation benchmarks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 24 citations worldwide. Full citation record

  1. GVCCS: A Dataset for Contrail Identification and Tracking on Visible Whole Sky Camera Sequences

    cs.CV 2025-07 conditional novelty 7.0 of 10

    GVCCS is the first open dataset of ground-based visible all-sky camera video with instance-level contrail masks, temporal tracking, and flight IDs, plus Mask2Former baselines.

  2. Neural Collapse by Design: Learning Class Prototypes on the Hypersphere

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Supervised classification reaches neural collapse by design via normalized prototype losses on the hypersphere, outperforming CE and SCL on ImageNet-1K and other benchmarks with faster convergence and better transfer.

  3. MapRF: Weakly Supervised Online HD Map Construction via NeRF-Guided Self-Training

    cs.CV 2025-11 unverdicted novelty 6.0 of 10

    MapRF reaches about 75% of fully supervised HD map accuracy on Argoverse 2 and nuScenes by generating view-consistent pseudo labels via a NeRF conditioned on map predictions and refining them with Map-to-Ray Matching ...

  4. gen2seg: Generative Models Enable Generalizable Instance Segmentation

    cs.CV 2025-05 unverdicted novelty 6.0 of 10

    Finetuning generative models on limited instance segmentation data produces zero-shot generalization to unseen object categories and styles, matching or exceeding supervised baselines like SAM on ambiguous boundaries.

  5. Open-Set LiDAR Panoptic Segmentation Guided by Uncertainty-Aware Learning

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Uncertainty-guided LiDAR panoptic segmentation (ULOPS) uses evidential learning and three uncertainty losses to segment unknown objects, outperforming prior open-set baselines on KITTI-360 and nuScenes.

  6. Evaluation of Embedding-Based and Generative Methods for LLM-Driven Document Classification: Opportunities and Challenges

    cs.IR 2026-04 conditional novelty 4.0 of 10

    Qwen2.5-VL with CoT prompting reaches 82% zero-shot accuracy on geoscience document classification, beating the best multimodal embedding model (QQMM) at 63%.

Pith tools