REVIEW 6 cited by
SPair-71k: A Large-scale Benchmark for Semantic Correspondence
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Establishing visual correspondences under large intra-class variations, which is often referred to as semantic correspondence or semantic matching, remains a challenging problem in computer vision. Despite its significance, however, most of the datasets for semantic correspondence are limited to a small amount of image pairs with similar viewpoints and scales. In this paper, we present a new large-scale benchmark dataset of semantically paired images, SPair-71k, which contains 70,958 image pairs with diverse variations in viewpoint and scale. Compared to previous datasets, it is significantly larger in number and contains more accurate and richer annotations. We believe this dataset will provide a reliable testbed to study the problem of semantic correspondence and will help to advance research in this area. We provide the results of recent methods on our new dataset as baselines for further research. Our benchmark is available online at http://cvlab.postech.ac.kr/research/SPair-71k/.
Forward citations
Cited by 6 Pith papers
-
Weakly-Supervised Learning of Dense Functional Correspondences
A weakly-supervised pipeline that distills VLM functional part knowledge and multi-view spatial structure into a model for dense cross-category functional correspondence, outperforming baselines on new synthetic and r...
-
Unsupervised Pixel-Level Semantic Left-Right Understanding of In-the-Wild Images
An unsupervised hybrid 3D-shape + image framework produces dense pixel-level semantic left-right labels for objects in wild images, outperforming prior feature-based baselines even on unseen categories.
-
Hidden in plain sight: VLMs overlook their visual representations
VLMs perform far worse than their own visual encoders on vision-centric tasks because the language model fails to use accessible visual information and instead follows its language priors.
-
Semantic Correspondence: Unified Benchmarking and a Strong Baseline
Fine-tuning the last layers of DINOv2, optionally with a lightweight cost aggregator, yields state-of-the-art semantic correspondence accuracy, and a new survey and benchmark consolidate the field's results.
-
Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation
Massive activations in DiTs are timestep-driven detail channels; suppressing them guides finer sampling and AdaLN-modulating them yields more discriminative dense features.
-
Towards Robust Semantic Correspondence: A Benchmark and Insights
The abstract promises an adverse-condition benchmark for semantic correspondence, yet the full text is a GRB magnetar analysis, so the claimed benchmark is unverifiable.
Discussion (0). Continue with ORCID to comment.