Pith. sign in

REVIEW 1 cited by

CostFormer:Cost Transformer for Cost Aggregation in Multi-view Stereo

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.10320 v1 pith:TVRDQ6EL submitted 2023-05-17 cs.CV

classification cs.CV
keywords costtransformeraggregationproposedcnnscostformermethodsmulti-view
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The core of Multi-view Stereo(MVS) is the matching process among reference and source pixels. Cost aggregation plays a significant role in this process, while previous methods focus on handling it via CNNs. This may inherit the natural limitation of CNNs that fail to discriminate repetitive or incorrect matches due to limited local receptive fields. To handle the issue, we aim to involve Transformer into cost aggregation. However, another problem may occur due to the quadratically growing computational complexity caused by Transformer, resulting in memory overflow and inference latency. In this paper, we overcome these limits with an efficient Transformer-based cost aggregation network, namely CostFormer. The Residual Depth-Aware Cost Transformer(RDACT) is proposed to aggregate long-range features on cost volume via self-attention mechanisms along the depth and spatial dimensions. Furthermore, Residual Regression Transformer(RRT) is proposed to enhance spatial attention. The proposed method is a universal plug-in to improve learning-based MVS methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    PartCATSeg improves open-vocabulary part segmentation by separating object- and part-level cost volumes, adding a compositional loss, and injecting DINO structural guidance, achieving over 10% h-IoU gains on three benchmarks.

Pith tools