REVIEW 3 cited by
Early Visual Concept Learning with Unsupervised Deep Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Automated discovery of early visual concepts from raw image data is a major open challenge in AI research. Addressing this problem, we propose an unsupervised approach for learning disentangled representations of the underlying factors of variation. We draw inspiration from neuroscience, and show how this can be achieved in an unsupervised generative model by applying the same learning pressures as have been suggested to act in the ventral visual stream in the brain. By enforcing redundancy reduction, encouraging statistical independence, and exposure to data with transform continuities analogous to those to which human infants are exposed, we obtain a variational autoencoder (VAE) framework capable of learning disentangled factors. Our approach makes few assumptions and works well across a wide variety of datasets. Furthermore, our solution has useful emergent properties, such as zero-shot inference and an intuitive understanding of "objectness".
Forward citations
Cited by 3 Pith papers
-
Lund jet images from generative and cycle-consistent adversarial networks
A least-squares GAN trained on Lund jet plane images reproduces the simulated jet substructure distribution to within a few percent, and a CycleGAN maps between jet categories such as parton-level vs detector-level or...
-
Geometric Disentanglement for Generative Latent Shape Models
An unsupervised VAE for 3D point clouds is structured so that latent variables separately control intrinsic shape and extrinsic pose, using Laplace-Beltrami spectra and hierarchical disentanglement penalties.
-
Autoencoding sensory substitution
Deep recurrent autoencoders convert images to shortened audio signals that incorporate hearing models, enabling above-chance hand posture discrimination and object reaching after a few hours of training instead of months.
Discussion (0). Continue with ORCID to comment.