REVIEW 3 cited by
CGI-Stereo: Accurate and Real-Time Stereo Matching via Context and Geometry Interaction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
In this paper, we propose CGI-Stereo, a novel neural network architecture that can concurrently achieve real-time performance, competitive accuracy, and strong generalization ability. The core of our CGI-Stereo is a Context and Geometry Fusion (CGF) block which adaptively fuses context and geometry information for more effective cost aggregation and meanwhile provides feedback to feature learning to guide more effective contextual feature extraction. The proposed CGF can be easily embedded into many existing stereo matching networks, such as PSMNet, GwcNet and ACVNet. The resulting networks show a significant improvement in accuracy. Specially, the model which incorporates our CGF with ACVNet ranks $1^{st}$ on the KITTI 2012 and 2015 leaderboards among all the published methods. We further propose an informative and concise cost volume, named Attention Feature Volume (AFV), which exploits a correlation volume as attention weights to filter a feature volume. Based on CGF and AFV, the proposed CGI-Stereo outperforms all other published real-time methods on KITTI benchmarks and shows better generalization ability than other real-time methods. Code is available at https://github.com/gangweiX/CGI-Stereo.
Forward citations
Cited by 3 Pith papers
-
ESMStereo: Enhanced ShuffleMixer Disparity Upsampling for Real-Time and Accurate Stereo Matching
A real-time stereo matching architecture whose Enhanced ShuffleMixer upsampler fuses disparity and image features to recover detail lost by compact cost volumes, reaching state-of-the-art speed-accuracy trade-offs.
-
EndoMUST: Monocular Depth Estimation for Robotic Endoscopy via End-to-end Multi-step Self-supervised Training
EndoMUST improves self-supervised monocular depth in endoscopy with a three-step training schedule that separates optical flow, intrinsic image decomposition, and DV-LoRA finetuning.
-
A Wavelet-based Stereo Matching Framework for Solving Frequency Convergence Inconsistency
A wavelet-based stereo matching framework with separate high and low frequency feature streams and a high-frequency preservation LSTM reports first-place results on KITTI 2012 and KITTI 2015.
Discussion (0). Continue with ORCID to comment.