REVIEW 6 cited by
CausalVAE: Structured Causal Disentanglement in Variational Autoencoder
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Learning disentanglement aims at finding a low dimensional representation which consists of multiple explanatory and generative factors of the observational data. The framework of variational autoencoder (VAE) is commonly used to disentangle independent factors from observations. However, in real scenarios, factors with semantics are not necessarily independent. Instead, there might be an underlying causal structure which renders these factors dependent. We thus propose a new VAE based framework named CausalVAE, which includes a Causal Layer to transform independent exogenous factors into causal endogenous ones that correspond to causally related concepts in data. We further analyze the model identifiabitily, showing that the proposed model learned from observations recovers the true one up to a certain degree by providing supervision signals (e.g. feature labels). Experiments are conducted on various datasets, including synthetic and real word benchmark CelebA. Results show that the causal representations learned by CausalVAE are semantically interpretable, and their causal relationship as a Directed Acyclic Graph (DAG) is identified with good accuracy. Furthermore, we demonstrate that the proposed CausalVAE model is able to generate counterfactual data through "do-operation" to the causal factors.
Forward citations
Cited by 6 Pith papers
-
Show Me Examples: Inferring Visual Concepts from Image Sets
Introduces VICIS task and training framework for inferring visual concepts from image sets, with experiments showing better accuracy, diversity, and generalization than standard VLMs on synthetic and ImageNet data.
-
PatchGen: Learning Soft Intra-Image Predictive Subsets for Visual Generalization
PatchGen learns a sample-dependent soft mask that selects label-predictive image patches, improving visual generalization across domain, category, and combined shifts.
-
TRACE-Seg3D: Counterfactual Context Auditing For Robust 3D Glioma Segmentation Under Institutional Shift
TRACE-Seg3D factorizes disease and imaging context representations in 3D MRI segmentation, auditing prediction stability under counterfactual context shifts to expose spurious failures missed by standard metrics.
-
Causally Steered Diffusion for Automated Video Counterfactual Generation
CSVC optimizes text prompts via vision-language-model feedback to steer frozen video diffusion editors toward causally consistent facial counterfactuals such as aging, gender change, beard addition, and baldness.
-
EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion
EchoVideo preserves identity in generated human videos by pre-fusing face, image, and text features, then training with stochastic shallow-feature dropout to reduce copy-paste artifacts.
-
Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement Learning
Contrastive Successor Features recover ground-truth RL states up to a linear map whenever the skill-conditioned transition differences follow a von Mises-Fisher distribution and policies are diverse.
Discussion (0). Continue with ORCID to comment.