REVIEW 5 cited by
Text2Layer: Layered Image Generation using Latent Diffusion Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Layer compositing is one of the most popular image editing workflows among both amateurs and professionals. Motivated by the success of diffusion models, we explore layer compositing from a layered image generation perspective. Instead of generating an image, we propose to generate background, foreground, layer mask, and the composed image simultaneously. To achieve layered image generation, we train an autoencoder that is able to reconstruct layered images and train diffusion models on the latent representation. One benefit of the proposed problem is to enable better compositing workflows in addition to the high-quality image output. Another benefit is producing higher-quality layer masks compared to masks produced by a separate step of image segmentation. Experimental results show that the proposed method is able to generate high-quality layered images and initiates a benchmark for future work.
Forward citations
Cited by 5 Pith papers
-
ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition
ReDesign turns screenshots into editable layer hierarchies by having a vision-language agent choose tools step by step and verify each split, beating layered-decomposition baselines on a new Figma edit-replay benchmark.
-
RaDL: Relation-aware Disentangled Learning for Multi-Instance Text-to-Image Generation
RaDL adds two attention modules to a Stable Diffusion based pipeline: Attribute Enhancement for per-instance attribute fidelity and Relation Attention that uses action verbs from the prompt to model inter-instance rel...
-
Rethinking Layered Graphic Design Generation with a Top-Down Approach
Accordion decomposes AI-generated raster designs into editable background, object, and vectorized text layers using a VLM-driven top-down planning pipeline.
-
PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models
A new open dataset and synthesis pipeline for high-quality multi-layer transparent images, plus a fine-tuned ART+ model that users preferred over the original ART in about 60 percent of comparisons.
-
LumiGen: An LVLM-Enhanced Iterative Framework for Fine-Grained Text-to-Image Generation
An LVLM-driven iterative text-to-image framework whose claimed performance scores are explicitly labeled fictitious, so no empirical result is established.
Discussion (0). Continue with ORCID to comment.