REVIEW 3 cited by
High Fidelity Visualization of What Your Self-Supervised Representation Knows About
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Discovering what is learned by neural networks remains a challenge. In self-supervised learning, classification is the most common task used to evaluate how good a representation is. However, relying only on such downstream task can limit our understanding of what information is retained in the representation of a given input. In this work, we showcase the use of a Representation Conditional Diffusion Model (RCDM) to visualize in data space the representations learned by self-supervised models. The use of RCDM is motivated by its ability to generate high-quality samples -- on par with state-of-the-art generative models -- while ensuring that the representations of those samples are faithful i.e. close to the one used for conditioning. By using RCDM to analyze self-supervised models, we are able to clearly show visually that i) SSL (backbone) representation are not invariant to the data augmentations they were trained with -- thus debunking an often restated but mistaken belief; ii) SSL post-projector embeddings appear indeed invariant to these data augmentation, along with many other data symmetries; iii) SSL representations appear more robust to small adversarial perturbation of their inputs than representations trained in a supervised manner; and iv) that SSL-trained representations exhibit an inherent structure that can be explored thanks to RCDM visualization and enables image manipulation.
Forward citations
Cited by 3 Pith papers
-
CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning
CheXWorld pre-trains a chest X-ray ViT with three world-modeling tasks (local anatomy, global layout, domain variation) and reports state-of-the-art transfer on eight benchmarks.
-
EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance
A motion-aware world model pretraining approach reduces echocardiography probe guidance error relative to existing visual backbones and guidance frameworks on a private clinical dataset.
-
Image Classification Using a Diffusion Model as a Pre-Training Model
Representation-conditioned diffusion pre-training improves hematoma classification accuracy by +6.15% and F1 by +13.60% over DINOv2 on a 179-image brain CT test set.
Discussion (0). Continue with ORCID to comment.