What Drives Compositional Generalization? The Importance of Continuous Training Objectives in Visual Generative Models

· 2025 · cs.CV · arXiv 2510.03075

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

open full Pith review browse 1 citing papers arXiv PDF

abstract

Compositional generalization, the ability to generate novel combinations of known concepts, is a key ingredient for visual generative models. Yet, not all mechanisms that enable or inhibit it are fully understood. In this work, we conduct a systematic study of how various design choices influence compositional generalization in image and video generation in a positive or negative way. Through controlled experiments, we identify two key factors: (i) whether the training objective operates on a discrete or continuous distribution, and (ii) to what extent conditioning provides information about the constituent concepts during training. Building on these insights, we show that relaxing the MaskGIT discrete loss with an auxiliary continuous JEPA-based objective can improve compositional performance in discrete models like MaskGIT.

representative citing papers

When Do Diffusion Models learn to Generate Multiple Objects?

cs.CV · 2026-04-30 · unverdicted · novelty 6.0

Using the mosaic controlled dataset framework, experiments show scene complexity dominates over concept imbalance in diffusion model failures for multi-object generation, with counting especially hard in low-data regimes and compositional generalization collapsing under held-out combinations.

citing papers explorer

Showing 1 of 1 citing paper after filters.

When Do Diffusion Models learn to Generate Multiple Objects? cs.CV · 2026-04-30 · unverdicted · none · ref 7 · internal anchor
Using the mosaic controlled dataset framework, experiments show scene complexity dominates over concept imbalance in diffusion model failures for multi-object generation, with counting especially hard in low-data regimes and compositional generalization collapsing under held-out combinations.

What Drives Compositional Generalization? The Importance of Continuous Training Objectives in Visual Generative Models

fields

years

verdicts

representative citing papers

citing papers explorer