Pith. sign in

REVIEW 4 cited by

Spatial Broadcast Decoder: A Simple Architecture for Learning Disentangled Representations in VAEs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1901.07017 v2 pith:SMAMKRI2 submitted 2019-01-21 cs.LG cs.CVstat.ML

classification cs.LGcs.CVstat.ML
keywords broadcastdecodervaesarchitecturelatentrepresentationsspatialdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a simple neural rendering architecture that helps variational autoencoders (VAEs) learn disentangled representations. Instead of the deconvolutional network typically used in the decoder of VAEs, we tile (broadcast) the latent vector across space, concatenate fixed X- and Y-"coordinate" channels, and apply a fully convolutional network with 1x1 stride. This provides an architectural prior for dissociating positional from non-positional features in the latent distribution of VAEs, yet without providing any explicit supervision to this effect. We show that this architecture, which we term the Spatial Broadcast decoder, improves disentangling, reconstruction accuracy, and generalization to held-out regions in data space. It provides a particularly dramatic benefit when applied to datasets with small objects. We also emphasize a method for visualizing learned latent spaces that helped us diagnose our models and may prove useful for others aiming to assess data representations. Finally, we show the Spatial Broadcast Decoder is complementary to state-of-the-art (SOTA) disentangling techniques and when incorporated improves their performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Visually Interpretable Oscillator Networks for Soft Continuum Robots from Video

    cs.RO 2025-11 unverdicted novelty 7.0 of 10

    An attention broadcast decoder localizes latent dynamics on robot images and, when coupled to oscillator networks, improves multi-step prediction error up to 5.7x on a two-segment soft continuum robot.

  2. CoLa: Chinese Character Decomposition with Compositional Latent Components

    cs.CV 2025-06 conditional novelty 7.0 of 10

    CoLa learns compositional latent components of Chinese characters via slot attention and matches them to printed templates, achieving strong zero-shot Chinese character recognition without human-defined decomposition.

  3. Farm-Level, In-Season Crop Identification for India

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A Google DeepMind team built a transformer-based system that maps 12 crops across India at farm level, in-season, with state-level area agreement of 94% (winter) and 75% (monsoon) against the 2023-24 census.

  4. An Interpretable Representation Learning Approach for Diffusion Tensor Imaging

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A 9x9 grid representation of DTI tract FA values, encoded by a beta-TCVAE with spatial broadcast decoder, yields improved downstream sex classification F1 and higher MIG than a 3D VAE baseline.

Pith tools