Pith. sign in

REVIEW 10 cited by

Transparent Image Layer Diffusion using Latent Transparency

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.17113 v4 pith:6QGYFPAJ submitted 2024-02-27 cs.CV cs.GR

classification cs.CVcs.GR
keywords latenttransparentdiffusionlayermodeltransparencyimagegeneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present LayerDiffuse, an approach enabling large-scale pretrained latent diffusion models to generate transparent images. The method allows generation of single transparent images or of multiple transparent layers. The method learns a "latent transparency" that encodes alpha channel transparency into the latent manifold of a pretrained latent diffusion model. It preserves the production-ready quality of the large diffusion model by regulating the added transparency as a latent offset with minimal changes to the original latent distribution of the pretrained model. In this way, any latent diffusion model can be converted into a transparent image generator by finetuning it with the adjusted latent space. We train the model with 1M transparent image layer pairs collected using a human-in-the-loop collection scheme. We show that latent transparency can be applied to different open source image generators, or be adapted to various conditional control systems to achieve applications like foreground/background-conditioned layer generation, joint layer generation, structural control of layer contents, etc. A user study finds that in most cases (97%) users prefer our natively generated transparent content over previous ad-hoc solutions such as generating and then matting. Users also report the quality of our generated transparent images is comparable to real commercial transparent assets like Adobe Stock.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LaRender: Training-Free Occlusion Control in Image Generation via Latent Rendering

    cs.CV 2025-08 conditional novelty 7.0 of 10

    LaRender replaces cross-attention layers in a pretrained diffusion model with a latent alpha-compositing operation that renders object features in occlusion order, giving training-free occlusion control.

  2. LayerFlow: A Unified Model for Layer-aware Video Generation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    LayerFlow is a unified diffusion-transformer model that generates transparent foreground, background, and blended video layers from per-layer prompts, and supports decomposition and conditioned generation in one framework.

  3. UniWorld-Design: From Pixel Generation to Layer-Native Design

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A two-model framework generates images as transparent layers and decomposes finished designs into ordered, complete semantic layers, outperforming prior decomposition models on per-layer fidelity and editability.

  4. Parallax Portrait Matting

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Parallax from a casually captured second view, fused asymmetrically (background-aligned pixels plus foreground-aligned cross-attention), yields finer portrait mattes and cleaner foreground colors than strong single-im...

  5. PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new open dataset and synthesis pipeline for high-quality multi-layer transparent images, plus a fine-tuned ART+ model that users preferred over the original ART in about 60 percent of comparisons.

  6. DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers

    cs.CV 2025-05 conditional novelty 6.0 of 10

    DiffDecompose recovers foreground and background layers from alpha-composited images using in-context diffusion with position encoding cloning, trained and evaluated on a new six-task synthetic dataset.

  7. Text-Conditioned Background Generation for Editable Multi-Layer Documents

    cs.CV 2025-12 conditional novelty 5.0 of 10

    A training-free system combines soft latent masking, WCAG-contrast-optimized semi-transparent text backings, and recursive LLM summaries to generate readable, style-consistent backgrounds for multi-page documents.

  8. All Stories Are One Story: Emotional Arc Guided Procedural Game Level Generation

    cs.AI 2025-08 conditional novelty 5.0 of 10

    A procedural game generation system that uses Rise/Fall emotional arcs to shape LLM-written branching stories and entity difficulty showed higher player enjoyment in a small ARPG user study.

  9. RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A decoupled-attention adapter transfers image-pair edits to new photos in diffusion transformers, trained with a new 218-task visual editing dataset.

  10. Hunyuan-Game: Industrial-grade Intelligent Game Creation Model

    cs.CV 2025-05 reject novelty 4.0 of 10

    Tencent's Hunyuan-Game applies diffusion transformers to game asset creation across nine image and video generation tasks, with self-reported gains that are partly contradicted by its own evaluation table.

Pith tools