REVIEW 4 cited by
Spatial Broadcast Decoder: A Simple Architecture for Learning Disentangled Representations in VAEs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present a simple neural rendering architecture that helps variational autoencoders (VAEs) learn disentangled representations. Instead of the deconvolutional network typically used in the decoder of VAEs, we tile (broadcast) the latent vector across space, concatenate fixed X- and Y-"coordinate" channels, and apply a fully convolutional network with 1x1 stride. This provides an architectural prior for dissociating positional from non-positional features in the latent distribution of VAEs, yet without providing any explicit supervision to this effect. We show that this architecture, which we term the Spatial Broadcast decoder, improves disentangling, reconstruction accuracy, and generalization to held-out regions in data space. It provides a particularly dramatic benefit when applied to datasets with small objects. We also emphasize a method for visualizing learned latent spaces that helped us diagnose our models and may prove useful for others aiming to assess data representations. Finally, we show the Spatial Broadcast Decoder is complementary to state-of-the-art (SOTA) disentangling techniques and when incorporated improves their performance.
Forward citations
Cited by 4 Pith papers
-
Learning Visually Interpretable Oscillator Networks for Soft Continuum Robots from Video
An attention broadcast decoder localizes latent dynamics on robot images and, when coupled to oscillator networks, improves multi-step prediction error up to 5.7x on a two-segment soft continuum robot.
-
CoLa: Chinese Character Decomposition with Compositional Latent Components
CoLa learns compositional latent components of Chinese characters via slot attention and matches them to printed templates, achieving strong zero-shot Chinese character recognition without human-defined decomposition.
-
Farm-Level, In-Season Crop Identification for India
A Google DeepMind team built a transformer-based system that maps 12 crops across India at farm level, in-season, with state-level area agreement of 94% (winter) and 75% (monsoon) against the 2023-24 census.
-
An Interpretable Representation Learning Approach for Diffusion Tensor Imaging
A 9x9 grid representation of DTI tract FA values, encoded by a beta-TCVAE with spatial broadcast decoder, yields improved downstream sex classification F1 and higher MIG than a 3D VAE baseline.
Discussion (0). Continue with ORCID to comment.