Pith. sign in

REVIEW 6 cited by

SegGPT: Segmenting Everything In Context

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.03284 v1 pith:WRIRWEJ7 submitted 2023-04-06 cs.CV

classification cs.CV
keywords segmentationseggpttaskscontextin-contextsegmentingdataeverything
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present SegGPT, a generalist model for segmenting everything in context. We unify various segmentation tasks into a generalist in-context learning framework that accommodates different kinds of segmentation data by transforming them into the same format of images. The training of SegGPT is formulated as an in-context coloring problem with random color mapping for each data sample. The objective is to accomplish diverse tasks according to the context, rather than relying on specific colors. After training, SegGPT can perform arbitrary segmentation tasks in images or videos via in-context inference, such as object instance, stuff, part, contour, and text. SegGPT is evaluated on a broad range of tasks, including few-shot semantic segmentation, video object segmentation, semantic segmentation, and panoptic segmentation. Our results show strong capabilities in segmenting in-domain and out-of-domain targets, either qualitatively or quantitatively.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A two-stage pure RL method with an information-gap global view and hierarchical grounding loss makes MLLMs truly rely on precise crops and sets SOTA on high-res VQA under tight token budgets.

  2. Stable Diffusion Models are Secretly Good at Visual In-Context Learning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A training-free attention recomputation inside Stable Diffusion self-attention enables visual in-context learning across six vision tasks.

  3. Decouple before Align: Visual Disentanglement Enhances Prompt Tuning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Decoupling images into foreground and background before aligning them with text improves CLIP prompt tuning on few-shot and generalization benchmarks.

  4. Generate Aligned Anomaly: Region-Guided Few-Shot Anomaly Image-Mask Pair Synthesis for Industrial Inspection

    cs.CV 2025-07 conditional novelty 6.0 of 10

    GAA synthesizes aligned anomaly image-mask pairs from few examples using decomposed concept embeddings and region-guided masks, improving downstream anomaly localization and classification on MVTec AD and LOCO.

  5. Is Visual in-Context Learning for Compositional Medical Tasks within Reach?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Training on synthetic compositional task sequences with sequence-level masking lets a transformer-based in-context learner follow multi-step medical imaging instructions on held-out images, but well below codebook upp...

  6. DOMR: Establishing Cross-View Segmentation via Dense Object Matching

    cs.CV 2025-08 conditional novelty 5.0 of 10

    DOMR jointly matches and refines multiple object masks across ego and exo views, reaching 49.7% and 55.2% mean IoU on Ego-Exo4D.

Pith tools