REVIEW 8 cited by
Disentangled Representation Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Disentangled Representation Learning (DRL) aims to learn a model capable of identifying and disentangling the underlying factors hidden in the observable data in representation form. The process of separating underlying factors of variation into variables with semantic meaning benefits in learning explainable representations of data, which imitates the meaningful understanding process of humans when observing an object or relation. As a general learning strategy, DRL has demonstrated its power in improving the model explainability, controlability, robustness, as well as generalization capacity in a wide range of scenarios such as computer vision, natural language processing, and data mining. In this article, we comprehensively investigate DRL from various aspects including motivations, definitions, methodologies, evaluations, applications, and model designs. We first present two well-recognized definitions, i.e., Intuitive Definition and Group Theory Definition for disentangled representation learning. We further categorize the methodologies for DRL into four groups from the following perspectives, the model type, representation structure, supervision signal, and independence assumption. We also analyze principles to design different DRL models that may benefit different tasks in practical applications. Finally, we point out challenges in DRL as well as potential research directions deserving future investigations. We believe this work may provide insights for promoting the DRL research in the community.
Forward citations
Cited by 8 Pith papers
-
The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail
Reliable concept presence in transformers is concentrated in the extreme high-activation tail of in-concept tokens; thresholding that tail improves concept detection and localization.
-
SEED: Speaker Embedding Enhancement Diffusion Model
SEED refines noisy speaker embeddings toward clean ones using a diffusion-style training objective, improving EER by up to 19.6% on a simulated mismatch benchmark with no speaker labels.
-
ConceptVAE: Self-Supervised Fine-Grained Concept Disentanglement from 2D Echocardiographies
ConceptVAE learns to discretize echocardiograms into fine-grained anatomical concepts and per-concept styles without labels, and reports gains over a VICReg baseline on retrieval, segmentation, and OOD detection.
-
Improving Generalization for AI-Synthesized Voice Detection
A disentanglement and sharpness-aware training framework improves cross-domain AI-synthesized voice detection by up to 7.59% EER over prior art.
-
URECA: The Chain of Two Minimum Set Cover Problems exists behind Adaptation to Shifts in Semantic Code Search
The paper derives (with a flawed Lebesgue-integral argument) that entropy minimization performs two-level set-cover clustering and introduces URECA, a union-find clustering loss that improves few-shot code-search adaptation.
-
Are Representation Disentanglement and Interpretability Linked in Recommendation Models? A Critical Review and Reproducibility Study
In recommender models, disentangled representations correlate with representation interpretability but not with recommendation effectiveness.
-
Imitation Learning Based on Disentangled Representation Learning of Behavioral Characteristics
A weakly-supervised CVAE with action chunking lets a robot change wiping speed online from instruction labels, but the same mechanism fails to disentangle wiping force and fails on spatial pick-and-place directives.
-
BOLDreams: Dreaming with pruned in-silico fMRI Encoding Models of the Visual Cortex
Pruned encoding models for visual cortex fMRI use as few as 1% of neural features per layer without losing accuracy, and different pretrained backbones predict the BOLD signal via distinct visual features.
Discussion (0). Continue with ORCID to comment.