Pith. sign in

REVIEW 5 cited by

Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.03163 v3 pith:ZW6XBW56 submitted 2024-03-05 cs.CL cs.CVcs.CY

classification cs.CLcs.CVcs.CY
keywords multimodalcodemetricsmodelswebpagesautomaticbenchmarkdesign2code
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative AI has made rapid advancements in recent years, achieving unprecedented capabilities in multimodal understanding and code generation. This can enable a new paradigm of front-end development in which multimodal large language models (MLLMs) directly convert visual designs into code implementations. In this work, we construct Design2Code - the first real-world benchmark for this task. Specifically, we manually curate 484 diverse real-world webpages as test cases and develop a set of automatic evaluation metrics to assess how well current multimodal LLMs can generate the code implementations that directly render into the given reference webpages, given the screenshots as input. We also complement automatic metrics with comprehensive human evaluations to validate the performance ranking. To rigorously benchmark MLLMs, we test various multimodal prompting methods on frontier models such as GPT-4o, GPT-4V, Gemini, and Claude. Our fine-grained break-down metrics indicate that models mostly lag in recalling visual elements from the input webpages and generating correct layout designs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation

    cs.HC 2026-05 conditional novelty 6.0 of 10

    MobileForge shows that current multimodal LLMs can compile multi-screen app projects from screenshots, but interactive navigation, visual fidelity, and code maintainability still fall short.

  2. GameDevBench: Evaluating Agentic Capabilities Through Game Development

    cs.AI 2026-02 conditional novelty 6.0 of 10

    A new 132-task Godot benchmark shows frontier AI agents solve only about 54.5% of game-development tasks, with visual feedback giving consistent but modest gains.

  3. Multilingual Multimodal Software Developer for Code Generation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A 7B vision-language model trained on synthetic diagram-to-code data outperforms several larger open-weight models on a new 10-language UML/flowchart code-generation benchmark.

  4. FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation

    cs.SE 2025-06 conditional novelty 6.0 of 10

    A new benchmark adds 148 interactive front-end development tasks with automated sandbox tests, reporting a 90.54% agreement rate with human evaluation across four LLMs.

  5. LOCOFY Large Design Models -- Design to code conversion solution

    cs.SE 2025-07 reject novelty 4.0 of 10

    A proprietary design-to-code pipeline is described with claimed high fidelity and LLM outperformance, but the evaluation is self-referential, unquantified, and unreproducible.

Pith tools