Pith. sign in

REVIEW 2 major objections 1 minor 1 cited by

Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learning

T0 review · 2 major / 1 minor · reviewed 2026-06-29 · grok-4.3

Pith's one-line read Stable-Layers fine-tunes layer decomposition models with VLM-scored reinforcement learning and no paired supervision.

desk verdict The two-stage VLM calibration gives a workable fix for reward variance in Flow-GRPO, but the paper shows no evidence that those scores track actual layer quality. read the letter →

arxiv 2605.30257 v1 pith:M3EGOZNB submitted 2026-05-28 cs.CV

classification cs.CV
keywords layerdecompositionreinforcementlearningvision-languagemodelsfine-tuningimageeditingGRPOCrellodataset
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper presents a reinforcement learning approach that starts from a pretrained layer decomposition model and improves it solely through feedback from a vision-language model. It solves the problem of narrow score ranges by using a two-stage pipeline: first scoring each candidate decomposition on five edit-centric criteria, then calibrating all candidates together in a grid view to restore variance. With this reward signal the method runs Flow-GRPO plus LoRA adaptation, sampling multiple decompositions per image and updating the policy from group-relative advantages. The result is decompositions that separate layers more cleanly, contain fewer blank or artifact-laden outputs, and show lower per-layer reconstruction error on the Crello dataset.

What carries the argument

The two-stage VLM evaluation pipeline that pairs structured per-sample scoring across five edit-centric criteria with a grid-based calibration step in which the VLM re-scores all candidates side-by-side.

What would settle it

Running the identical training loop but omitting the grid-based calibration step and checking whether within-group score variance collapses to the point that policy updates become negligible, or measuring whether the reported gains on layer separation and reconstruction error vanish on a fresh held-out image set.

Watch

Extended reading notes

Core claim

Starting from Qwen-Image-Layered, Stable-Layers applies Flow-GRPO with LoRA adaptation, sampling multiple candidate decompositions per image, scoring them with a VLM through a two-stage pipeline of per-sample criteria scoring followed by grid-based side-by-side calibration, and optimising the policy from group-relative advantages, yielding stronger layer separation, fewer blank or artifact-heavy layers, and lower per-layer reconstruction error on Crello compared with the base model.

Load-bearing premise

The two-stage VLM evaluation pipeline supplies a sufficiently reliable and high-variance reward signal that enables effective policy improvement via Flow-GRPO.

Editorial extensions

If this is right

  • Decompositions exhibit stronger layer separation than the base model.
  • The number of blank or artifact-heavy layers decreases.
  • Per-layer reconstruction error drops on the Crello dataset.
  • Fine-tuning becomes possible without any paired ground-truth decompositions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same calibrated VLM reward construction could be tested on other generative tasks where human preference data is expensive but visual quality is easy to judge.
  • If the calibration step proves essential, future VLM-as-reward pipelines may need an explicit group comparison stage rather than isolated scoring.
  • The method could be re-run with different base VLMs to test whether the quality of the reward model itself limits further gains.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper proposes Stable-Layers, a reinforcement learning framework for fine-tuning a pretrained image layer decomposition model (Qwen-Image-Layered) using only VLM feedback via Flow-GRPO with LoRA. It introduces a two-stage VLM scoring pipeline (per-sample scoring on five criteria followed by grid-based side-by-side calibration) to address score compression and low variance in rewards. The method is claimed to produce better layer decompositions with stronger separation, fewer artifacts, and lower reconstruction error on the Crello dataset compared to the base model.

Significance. If the VLM-based reward signal proves reliable, this approach could enable effective fine-tuning of decomposition models without requiring paired supervision data, which is often scarce. However, the current manuscript provides no quantitative results or validation of the reward, limiting the ability to assess its significance.

major comments (2)
  1. [Abstract] Abstract: the abstract states that Stable-Layers produces decompositions with stronger layer separation, fewer blank or artifact-heavy layers, and lower per-layer reconstruction error on the Crello dataset, but supplies no quantitative results, error bars, ablation details, or dataset statistics to support these claims.
  2. [Abstract] Abstract: the central claim relies on the two-stage VLM evaluation pipeline supplying a reliable reward signal, but no correlation analysis is provided between the VLM scores and human expert ratings or objective proxies such as layer-wise IoU against Crello ground-truth masks.
minor comments (1)
  1. The description of the five edit-centric criteria could be expanded for reproducibility.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for highlighting the need for stronger quantitative support in the abstract and validation of the VLM-based reward. We will revise the manuscript accordingly to address these points while preserving the core contributions of the two-stage scoring pipeline and Flow-GRPO fine-tuning approach.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the abstract states that Stable-Layers produces decompositions with stronger layer separation, fewer blank or artifact-heavy layers, and lower per-layer reconstruction error on the Crello dataset, but supplies no quantitative results, error bars, ablation details, or dataset statistics to support these claims.

    Authors: We agree the abstract should be more self-contained with quantitative backing. In the revision we will insert concrete metrics drawn from the experimental section (e.g., mean per-layer reconstruction error reduction, layer-separation scores, and Crello dataset statistics) together with error bars from repeated runs and a brief reference to the ablation studies already present in the main text. revision: yes

  2. Referee: [Abstract] Abstract: the central claim relies on the two-stage VLM evaluation pipeline supplying a reliable reward signal, but no correlation analysis is provided between the VLM scores and human expert ratings or objective proxies such as layer-wise IoU against Crello ground-truth masks.

    Authors: We accept that explicit validation of the reward signal strengthens the central claim. The revised manuscript will include a new subsection reporting layer-wise IoU correlations against Crello ground-truth masks and, where feasible, a small-scale human rating study. If resource constraints limit the human study, we will clearly state this as a limitation while still providing the IoU analysis. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity detected in derivation chain

full rationale

The paper's core method applies Flow-GRPO to a pretrained layer decomposition model using reward signals generated by an external VLM through a two-stage per-sample and grid-calibration pipeline. This reward is independent of any quantities defined or fitted inside the decomposition model itself, and the reported improvements (layer separation, reconstruction error on Crello) are measured against the base model on held-out data rather than being tautological to the training objective. No self-definitional loops, fitted inputs renamed as predictions, or load-bearing self-citations appear in the derivation; the approach is externally grounded in VLM feedback and dataset evaluation.

Assumptions & free parameters 2 free parameters · 1 assumptions · 0 invented entities

Based solely on the abstract; full implementation details unavailable. The primary unverified premise is that VLM judgments on the chosen criteria correlate with human-preferred layer quality.

free parameters (2)
  • five edit-centric criteria
    Chosen by authors to structure VLM scoring; exact definitions and weighting not specified in abstract.
  • grid calibration parameters
    Number of candidates per group and grid layout are design choices that affect score variance.
assumptions (1)
  • domain assumption Vision-language models can produce consistent and informative scores for image layer decompositions when given structured per-sample criteria plus side-by-side comparison.
    This assumption underpins the entire reward signal and is stated as the key challenge the paper addresses.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learning." pith.science (2026). https://pith.science/paper/M3EGOZNB

@misc{pith2026260530257,
  author       = {Pith},
  title        = {Pith review of: Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/M3EGOZNB}},
  note         = {Machine review of arXiv:2605.30257}
}
read the original abstract

We present Stable-Layers, a reinforcement learning framework that eliminates the need for paired supervision by fine-tuning a pretrained layer decomposition model using only feedback from a vision-language model (VLM). Starting from Qwen-Image-Layered, we apply Flow-GRPO with LoRA adaptation, sampling multiple candidate decompositions per image, scoring them with a VLM, and optimising the policy from group-relative advantages. The key challenge lies in designing a reliable reward signal: VLMs scoring samples in isolation tend to compress their judgements into a narrow band, leaving GRPO with little within-group variance to learn from. We address this with a two-stage evaluation pipeline that pairs structured per-sample scoring across five edit-centric criteria with a grid-based calibration step in which the VLM re-scores all candidates side-by-side. Stable-Layers produces decompositions with stronger layer separation, fewer blank or artifact-heavy layers, and lower per-layer reconstruction error on the Crello dataset compared to the base model.

Figures

Figures reproduced from arXiv: 2605.30257 by the authors.

Figure 1
Figure 1. Stable-Layers. We finetune a layer decomposition model using Flow-GRPO and a VLM judge, and improve layerization without relying on paired data. The resulting layers have improved consistency, separation and handle in-painting of occluded areas better. Abstract We present Stable-Layers, a reinforcement learning framework that eliminates the need for paired supervision by fine-tuning a pretrained layer decomposition … view at source ↗
Figure 2
Figure 2. Stable-Layers training pipeline. Sample G candidates, score with the two-phase VLM reward, replay with GRPO updates to LoRA parameters. variance, preventing the clipped surrogate from constraining overconfident positive-advantage updates. GRPO-Guard addresses this with two corrections: (i) RatioNorm, which standardizes log ρg per denoising step so that the ratio distribution is centered near 1 with uniform variance … view at source ↗
Figure 4
Figure 4. Held-out evaluation metrics. Three auto￾mated metrics on 480 LAION-Aesthetics [22] images across training. Top: bad layers per decomposition (blank + glaze; lower is better) fall from ∼1.65 to ∼0.4. Middle: feature distribution evenness (higher is better) rises from ∼0.53 to ∼0.73. Bottom: layer 0 inpainting quality (higher is better) rises from ∼0.38 to ∼0.62. Bands show ±1σ across the set. 6.2 Qualitative Results … view at source ↗
Figures from the paper (4 more)
Figure 5
Figure 5. Figure 5: Qualitative comparison on held-out images. Base model (Qwen Image Layered, top of each pair) vs. Stable-Layers-fine-tuned model (Stable-Layers, bottom). Columns show the input, the composite, and individual layers on white backgrounds. The fine-tuned model produces pla…
Figure 6
Figure 6. Figure 6: Extended qualitative gallery. Layer decompositions from the Stable-Layers-fine-tuned model on a diverse set of held-out inputs. Each row shows the reconstructed composite and the individual layers composited onto white backgrounds. Examples illustrate consistent behavi…
Figure 7
Figure 7. Figure 7: Calibration ablation (additional Layer 0 metrics). Layer 0 combined quality (left) and edge-density sharpness (right) over training, comparing the full two-phase reward with grid calibration against Phase 1 individual scoring alone. Both metrics show the calibrated run…
Figure 8
Figure 8. Figure 8: Mean VLM reward during training. Phase 2 calibrated reward per step (light) and rolling average (dark). The mean reward rises over the first ∼100 steps as the policy eliminates the worst failure modes, then plateaus around 0.83–0.85 with high per-step variance. Under G…

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. UniWorld-Design: From Pixel Generation to Layer-Native Design

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A two-model framework generates images as transparent layers and decomposes finished designs into ordered, complete semantic layers, outperforming prior decomposition models on per-layer fidelity and editability.

Reference graph

Works this paper leans on

45 extracted references · 8 canonical work pages · cited by 1 Pith paper

  1. [1]

    Training diffusion models with reinforcement learning

    Kevin Black, Michael Janner, Yilun Du, Ilya Kostrikov, and Sergey Levine. Training diffusion models with reinforcement learning. InInternational Conference on Learning Representations (ICLR), 2024

  2. [2]

    Referring layer decomposition.arXiv preprint arXiv:2602.19358, 2026

    Fangyi Chen, Yaojie Shen, Lu Xu, Ye Yuan, Shu Zhang, Yulei Niu, and Longyin Wen. Referring layer decomposition.arXiv preprint arXiv:2602.19358, 2026

  3. [3]

    From inpainting to layer decomposition: Repurposing generative inpainting models for image layer decomposition.arXiv preprint arXiv:2511.20996, 2025

    Jingxi Chen, Yixiao Zhang, Xiaoye Qian, Zongxia Li, Cornelia Fermuller, Caren Chen, and Yiannis Aloimonos. From inpainting to layer decomposition: Repurposing generative inpainting models for image layer decomposition.arXiv preprint arXiv:2511.20996, 2025

  4. [4]

    Topreward: Token probabilities as hidden zero-shot rewards for robotics.arXiv preprint arXiv:2602.19313, 2026

    Shirui Chen, Cole Harrison, Ying-Chun Lee, Angela Jin Yang, Zhongzheng Ren, Lillian J Ratliff, Jiafei Duan, Dieter Fox, and Ranjay Krishna. Topreward: Token probabilities as hidden zero-shot rewards for robotics.arXiv preprint arXiv:2602.19313, 2026

  5. [5]

    MJ-bench: Is your multimodal reward model really a good judge for text-to-image generation?arXiv preprint, 2024

    Zhaorun Chen, Yichao Du, Zichen Wen, Yiyang Zhou, Chenhang Cui, Zhenzhen Weng, Haoqin Tu, Chaoqi Wang, Zhe Tong, Qing Huang, Canyu Chen, Qinghao Ye, Zhihong Zhu, Yuqing Zhang, Jiawei Zhou, Zhuokai Zhao, Rafael Rafailov, Chelsea Finn, and Huaxiu Yao. MJ-bench: Is your multimodal reward model really a good judge for text-to-image generation?arXiv preprint, 2024

  6. [6]

    Kevin Clark, Paul Vicol, Kevin Swersky, and David J. Fleet. Directly fine-tuning diffusion models on differentiable rewards. InInternational Conference on Learning Representations (ICLR), 2024

  7. [7]

    LayerFusion: Harmonized multi-layer text-to-image generation with generative priors.arXiv preprint, 2024

    Yusuf Dalva, Yijun Li, Qing Liu, Nanxuan Zhao, Jianming Zhang, Zhe Lin, and Pinar Yanardag. LayerFusion: Harmonized multi-layer text-to-image generation with generative priors.arXiv preprint, 2024

  8. [8]

    DPOK: Reinforcement learning for fine-tuning text-to-image diffusion models

    Ying Fan, Olivia Watkins, Yuqing Du, Hao Liu, Moonkyung Ryu, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, Kangwook Lee, and Kimin Lee. DPOK: Reinforcement learning for fine-tuning text-to-image diffusion models. InNeural Information Processing Systems (NeurIPS), 2023

Show all 45 references
  1. [9]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models.arXiv preprint, 2021

  2. [10]

    PSDiffusion: Harmonized multi-layer image generation via layout and appearance alignment

    Dingbang Huang, Wenbo Li, Yifei Zhao, Xinyu Pan, Yanhong Zeng, and Bo Dai. PSDiffusion: Harmonized multi-layer image generation via layout and appearance alignment. InWinter Conference on Applications of Computer Vision (WACV), 2026

  3. [11]

    Dreamlayer: Simultaneous multi-layer generation via diffusion model

    Junjia Huang, Pengxiang Yan, Jinhang Cai, Jiyang Liu, Zhao Wang, Yitong Wang, Xinglong Wu, and Guanbin Li. Dreamlayer: Simultaneous multi-layer generation via diffusion model. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3357–3366, 2025

  4. [12]

    LayerDiff: Exploring text-guided multi-layered composable image synthesis via layer-collaborative diffusion model

    Runhui Huang et al. LayerDiff: Exploring text-guided multi-layered composable image synthesis via layer-collaborative diffusion model. InEuropean Conference on Computer Vision (ECCV), 2024

  5. [13]

    Layeringdiff: Layered image synthesis via generation, then disassembly with generative knowledge.arXiv preprint arXiv:2501.01197, 2025

    Kyoungkook Kang, Gyujin Sim, Geonung Kim, Donguk Kim, Seungho Nam, and Sunghyun Cho. Layeringdiff: Layered image synthesis via generation, then disassembly with generative knowledge.arXiv preprint arXiv:2501.01197, 2025

  6. [14]

    Self-rewarding vision-language model via reasoning decomposition.arXiv preprint arXiv:2508.19652, 2025

    Zongxia Li, Wenhao Yu, Chengsong Huang, Rui Liu, Zhenwen Liang, Fuxiao Liu, Jingxi Che, Dian Yu, Jordan Boyd-Graber, Haitao Mi, et al. Self-rewarding vision-language model via reasoning decomposition.arXiv preprint arXiv:2508.19652, 2025

  7. [15]

    See-through: Single-image layer decomposition for anime characters.arXiv preprint arXiv:2602.03749, 2026

    Jian Lin, Chengze Li, Haoyun Qin, Kwun Wang Chan, Yanghua Jin, Hanyuan Liu, Stephen Chun Wang Choy, and Xueting Liu. See-through: Single-image layer decomposition for anime characters.arXiv preprint arXiv:2602.03749, 2026

  8. [16]

    Flow-grpo: Training flow matching models via online rl.arXiv preprint arXiv:2505.05470, 2025

    Jie Liu, Gongye Liu, Jiajun Liang, Yangguang Li, Jiaheng Liu, Xintao Wang, Pengfei Wan, Di Zhang, and Wanli Ouyang. Flow-grpo: Training flow matching models via online rl.arXiv preprint arXiv:2505.05470, 2025. 10

  9. [17]

    Rectified flow: A marginal preserving approach to optimal transport.arXiv preprint, 2022

    Qiang Liu. Rectified flow: A marginal preserving approach to optimal transport.arXiv preprint, 2022

  10. [18]

    Flow straight and fast: Learning to generate and transfer data with rectified flow.arXiv preprint, 2022

    Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow.arXiv preprint, 2022

  11. [19]

    Controllable layer decomposition for reversible multi-layer image generation.arXiv preprint, 2025

    Zihao Liu, Zunnan Xu, Shi Shu, Jun Zhou, Ruicheng Zhang, Zhenchao Tang, and Xiu Li. Controllable layer decomposition for reversible multi-layer image generation.arXiv preprint, 2025

  12. [20]

    Fine-T2I: An open, large-scale, and diverse dataset for high-quality T2I fine-tuning.arXiv preprint, 2026

    Xu Ma, Yitian Zhang, Qihua Dong, and Yun Fu. Fine-T2I: An open, large-scale, and diverse dataset for high-quality T2I fine-tuning.arXiv preprint, 2026

  13. [21]

    Aligning text-to- image diffusion models with reward backpropagation.arXiv preprint, 2023

    Mihir Prabhudesai, Anirudh Goyal, Deepak Pathak, and Katerina Fragkiadaki. Aligning text-to- image diffusion models with reward backpropagation.arXiv preprint, 2023

  14. [22]

    LAION- 5B: An open large-scale dataset for training next generation image-text models.Neural Infor- mation Processing Systems (NeurIPS), 35, 2022

    Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. LAION- 5B: An open large-scale dataset for training next generation image-text models.Neural Infor- mation Proce...

  15. [23]

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y . K. Li, Y . Wu, and Daya Guo. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint, 2024

  16. [24]

    LayerD: Decomposing raster graphic designs into layers

    Tomoyuki Suzuki, Kang-Jun Liu, Naoto Inoue, and Kota Yamaguchi. LayerD: Decomposing raster graphic designs into layers. InInternational Conference on Computer Vision (ICCV), 2025

  17. [25]

    Diffusion model alignment using direct preference optimization

    Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purber, Stefano Ermon, Caiming Xiong, Shafiq Joty, and Nikhil Naik. Diffusion model alignment using direct preference optimization. InConference on Computer Vision and Pattern Recognition (CVPR), 2024

  18. [26]

    Grpo-guard: Mitigating implicit over-optimization in flow matching via regulated clipping.arXiv preprint, 2025

    Jing Wang, Jiajun Liang, Jie Liu, Henglin Liu, Gongye Liu, Jun Zheng, Wanyuan Pang, Ao Ma, Zhenyu Xie, Xintao Wang, et al. Grpo-guard: Mitigating implicit over-optimization in flow matching via regulated clipping.arXiv preprint, 2025

  19. [27]

    Self-evolving vision-language models for image quality assessment via voting and ranking.arXiv preprint, 2025

    Wen Wen, Tianwu Zhi, Kanglong Fan, Yang Li, Xinge Peng, Yabin Zhang, Yiting Liao, Junlin Li, and Li Zhang. Self-evolving vision-language models for image quality assessment via voting and ranking.arXiv preprint, 2025

  20. [28]

    Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis.arXiv preprint, 2023

    Xiaoshi Wu, Yiming Hao, Keqiang Sun, Yixiong Chen, Feng Zhu, Rui Zhao, and Hongsheng Li. Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis.arXiv preprint, 2023

  21. [29]

    Visionreward: Fine-grained multi-dimensional human preference learning for image and video generation

    Jiazheng Xu, Yu Huang, Jiale Cheng, Yuanming Yang, Jiajun Xu, Yuan Wang, Wenbo Duan, Shen Yang, Qunlin Jin, Shurun Li, et al. Visionreward: Fine-grained multi-dimensional human preference learning for image and video generation. InAAAI Conference on Artificial Intelligence (AA...

  22. [30]

    ImageReward: Learning and evaluating human preferences for text-to-image generation

    Jiazheng Xu et al. ImageReward: Learning and evaluating human preferences for text-to-image generation. InNeural Information Processing Systems (NeurIPS), 2023

  23. [31]

    DanceGRPO: Unleashing GRPO on visual generation.arXiv preprint, 2025

    Zeyue Xue, Jie Wu, Yu Gao, Fangyuan Kong, Lingting Zhu, Mengzhao Chen, Zhiheng Liu, Wei Liu, Qiushan Guo, Weilin Huang, and Ping Luo. DanceGRPO: Unleashing GRPO on visual generation.arXiv preprint, 2025

  24. [32]

    CanvasV AE: Learning to generate vector graphic documents.International Conference on Computer Vision (ICCV), 2021

    Kota Yamaguchi. CanvasV AE: Learning to generate vector graphic documents.International Conference on Computer Vision (ICCV), 2021

  25. [33]

    Generative image layer decomposition with visual effects

    Jinrui Yang, Qing Liu, Yijun Li, Soo Ye Kim, Daniil Pakhomov, Mengwei Ren, Jianming Zhang, Zhe Lin, Cihang Xie, and Yuyin Zhou. Generative image layer decomposition with visual effects. InConference on Computer Vision and Pattern Recognition (CVPR), pages 7643–7653, 2025

  26. [34]

    Controllable layered image generation for real-world editing.arXiv preprint, 2026

    Jinrui Yang, Qing Liu, Yijun Li, Mengwei Ren, Letian Zhang, Zhe Lin, Cihang Xie, and Yuyin Zhou. Controllable layered image generation for real-world editing.arXiv preprint, 2026

  27. [35]

    Qwen-image-layered: Towards inherent editability via layer decomposition.arXiv preprint arXiv:2512.15603, 2025

    Shengming Yin, Zekai Zhang, Zecheng Tang, Kaiyuan Gao, Xiao Xu, Kun Yan, Jiahao Li, Yilei Chen, Yuxiang Chen, Heung-Yeung Shum, et al. Qwen-image-layered: Towards inherent editability via layer decomposition.arXiv preprint arXiv:2512.15603, 2025. 11

  28. [36]

    Transparent image layer diffusion using latent trans- parency.ACM Transactions on Graphics (TOG), 2024

    Lvmin Zhang and Maneesh Agrawala. Transparent image layer diffusion using latent trans- parency.ACM Transactions on Graphics (TOG), 2024

  29. [37]

    Direct preference optimization of video large multimodal models from language model reward

    Ruohong Zhang, Liangke Gui, Zhiqing Sun, Yihao Feng, Keyang Xu, Yuanhan Zhang, Di Fu, Chunyuan Li, Alexander G Hauptmann, Yonatan Bisk, et al. Direct preference optimization of video large multimodal models from language model reward. InProceedings of the 2025 Conference of th...

  30. [38]

    Workflow-aware structured layer decomposition for illustration production.arXiv preprint, 2026

    Tianyu Zhang, Dongchi Li, Keiichi Sawada, and Haoran Xie. Workflow-aware structured layer decomposition for illustration production.arXiv preprint, 2026. 12 A Algorithm Pseudocode Algorithm 1Stable-Layers Training Step (adapted from Flow-GRPO with GRPO-Guard stabilisation) Req...

  31. [39]

    a person, a car, a tree)

    semantic_separation (0-5): Each foreground layer should contain ONE distinct, complete object or semantic element (e.g. a person, a car, a tree). Score 0 if a single object is arbitrarily split across multiple layers or if layers contain random crops/slices of the scene rather...

  32. [40]

    Score 0 if layers show a semi-transparent haze, ghosting, colour bleed, or a milky/glazed wash over areas that should be fully transparent

    alpha_cleanliness (0-5): Foreground layers should have crisp, binary-like alpha with clean edges. Score 0 if layers show a semi-transparent haze, ghosting, colour bleed, or a milky/glazed wash over areas that should be fully transparent. Transparent regions must be FULLY trans...

  33. [41]

    Score 0 if the background is blurry, has obvious holes, smeared patches, or copy-paste artifacts where foreground objects were removed

    background_inpainting (0-5): Layer 0 (the background) should look like a plausible complete scene with foreground objects removed and their regions filled in convincingly. Score 0 if the background is blurry, has obvious holes, smeared patches, or copy-paste artifacts where fo...

  34. [42]

    Score 0 if most content is crammed into one layer while others are blank or near-empty

    feature_distribution (0-5): Visual content should be meaningfully spread across layers. Score 0 if most content is crammed into one layer while others are blank or near-empty. Score 5 if layers have a balanced, meaningful distribution of the scene’s content

  35. [43]

    semantic_separation

    content_validity (0-5): Penalize blank, empty, or noise-only layers. Score 0 if most layers are blank or contain only noise/blur. Score 5 if all layers have clear, recognizable content. - total (0-25): Sum of all five scores. Return ONLY valid JSON: {"semantic_separation":X, "...

  36. [44]

    Limitations

    found that this asymmetry did not degrade final sample quality while substantially reducing training cost. Trajectory replay.For each stored SDE step i, the current policy’s transition meanµθ i is recomputed via a forward pass through the LoRA-adapted transformer with gradient...

  37. [45]

    • Depending on the country in which research is conducted, IRB approval (or equivalent) may be required for any human subjects research

    Institutional review board (IRB) approvals or equivalent for research with human subjects Question: Does the paper describe potential risks incurred by study participants, whether such risks were disclosed to the subjects, and whether Institutional Review Board (IRB) approvals...

Pith tools

Reviewed June 29, 2026 · model on record in the stance chip above.