Pith. sign in

REVIEW 4 cited by

Graphic Design with Large Multimodal Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.14368 v1 pith:QLZWHCPB submitted 2024-04-22 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords designgraphicgenerationgraphistlayoutelementsfieldlarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the field of graphic design, automating the integration of design elements into a cohesive multi-layered artwork not only boosts productivity but also paves the way for the democratization of graphic design. One existing practice is Graphic Layout Generation (GLG), which aims to layout sequential design elements. It has been constrained by the necessity for a predefined correct sequence of layers, thus limiting creative potential and increasing user workload. In this paper, we present Hierarchical Layout Generation (HLG) as a more flexible and pragmatic setup, which creates graphic composition from unordered sets of design elements. To tackle the HLG task, we introduce Graphist, the first layout generation model based on large multimodal models. Graphist efficiently reframes the HLG as a sequence generation problem, utilizing RGB-A images as input, outputs a JSON draft protocol, indicating the coordinates, size, and order of each element. We develop new evaluation metrics for HLG. Graphist outperforms prior arts and establishes a strong baseline for this field. Project homepage: https://github.com/graphic-design-ai/graphist

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. IGD: Instructional Graphic Design with Multimodal Layer Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    IGD generates editable multi-layer graphic designs (posters, slides, stickers) from text instructions using an MLLM for layout and a diffusion model for image assets.

  2. Rethinking Layered Graphic Design Generation with a Top-Down Approach

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Accordion decomposes AI-generated raster designs into editable background, object, and vectorized text layers using a VLM-driven top-down planning pipeline.

  3. PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new open dataset and synthesis pipeline for high-quality multi-layer transparent images, plus a fine-tuned ART+ model that users preferred over the original ART in about 60 percent of comparisons.

  4. Aggregated Structural Representation with Large Language Models for Human-Centric Layout Generation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    ASR replaces the vision encoder of a multimodal LLM with graph-derived structural features to generate UI layouts, reporting better overlap and relation metrics than four prior methods.

Pith tools