Pith. sign in

REVIEW 1 cited by

Semantic Segmentation on VSPW Dataset through Aggregation of Transformer Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.01316 v1 pith:GRCMJNFO submitted 2021-09-03 cs.CV

classification cs.CV
keywords videoparsingscenesegmentationsemantictransformeraggregationchallenge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Semantic segmentation is an important task in computer vision, from which some important usage scenarios are derived, such as autonomous driving, scene parsing, etc. Due to the emphasis on the task of video semantic segmentation, we participated in this competition. In this report, we briefly introduce the solutions of team 'BetterThing' for the ICCV2021 - Video Scene Parsing in the Wild Challenge. Transformer is used as the backbone for extracting video frame features, and the final result is the aggregation of the output of two Transformer models, SWIN and VOLO. This solution achieves 57.3% mIoU, which is ranked 3rd place in the Video Scene Parsing in the Wild Challenge.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action Recognition

    cs.CV 2024-11 conditional novelty 6.0 of 10

    TAMT pre-trains on a source video dataset and fine-tunes small temporal-aware adapters plus a global moment-pooling head on the target, reporting new state-of-the-art cross-domain few-shot action recognition results.

Pith tools