Pith. sign in

REVIEW 6 cited by

PosterLLaVa: Constructing a Unified Multi-modal Layout Generator with LLM

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.02884 v3 pith:7ZWN23BI submitted 2024-06-05 cs.CV

classification cs.CV
keywords designgenerationlayoutgraphicmulti-modalautomateddatasetsmethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Layout generation is the keystone in achieving automated graphic design, requiring arranging the position and size of various multi-modal design elements in a visually pleasing and constraint-following manner. Previous approaches are either inefficient for large-scale applications or lack flexibility for varying design requirements. Our research introduces a unified framework for automated graphic layout generation, leveraging the multi-modal large language model (MLLM) to accommodate diverse design tasks. In contrast, our data-driven method employs structured text (JSON format) and visual instruction tuning to generate layouts under specific visual and textual constraints, including user-defined natural language specifications. We conducted extensive experiments and achieved state-of-the-art (SOTA) performance on public multi-modal layout generation benchmarks, demonstrating the effectiveness of our method. Moreover, recognizing existing datasets' limitations in capturing the complexity of real-world graphic designs, we propose two new datasets for much more challenging tasks (user-constrained generation and complicated poster), further validating our model's utility in real-life settings. Marking by its superior accessibility and adaptability, this approach further automates large-scale graphic design tasks. Finally, we develop an automated text-to-poster system that generates editable SVG posters based on users' design intentions, bridging the gap between layout generation and real-world graphic design applications. This system integrates our proposed layout generation method as the core component, demonstrating its effectiveness in practical scenarios. The code and datasets are open-sourced on https://github.com/posterllava/PosterLLaVA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SciFig: Towards Automating Editable Figure Generation for Scientific Papers

    cs.AI 2026-01 conditional novelty 6.0 of 10

    SciFig automatically generates editable methodology figures from scientific text and claims state-of-the-art quality on its own SciFig-Eval rubric-based benchmark.

  2. SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based Reflection

    cs.AI 2025-09 conditional novelty 6.0 of 10

    SheetDesigner uses zero-shot multimodal LLMs with rule- and vision-based reflection to generate spreadsheet layouts, and claims a 22.6% gain over baselines on a new seven-criterion benchmark.

  3. IGD: Instructional Graphic Design with Multimodal Layer Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    IGD generates editable multi-layer graphic designs (posters, slides, stickers) from text instructions using an MLLM for layout and a diffusion model for image assets.

  4. CAL-RAG: Retrieval-Augmented Multi-Agent Generation for Content-Aware Layout Design

    cs.IR 2025-06 reject novelty 5.0 of 10

    CAL-RAG reports state-of-the-art layout metrics on PKU PosterLayout by iteratively refining layouts with an agentic loop, but the perfect scores likely reflect direct optimization of the reported metrics.

  5. PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework

    cs.CV 2025-06 conditional novelty 5.0 of 10

    PosterCraft improves text-to-poster generation by cascading four stages of training (text rendering, region-weighted fine-tuning, preference optimization, and vision-language feedback), outperforming open-source basel...

  6. Aggregated Structural Representation with Large Language Models for Human-Centric Layout Generation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    ASR replaces the vision encoder of a multimodal LLM with graph-derived structural features to generate UI layouts, reporting better overlap and relation metrics than four prior methods.

Pith tools