TaoD2C-Bench: Benchmarking MLLMs for Industrial UI Code Generation Beyond Visual Fidelity
Chengwei Shi, Yunnong Chen, Tingting Zhou, Qiang Lu, Shiyu Yue, Xinyuan Hu, Jianfang Ru, Liuqing Chen
Abstract
A key challenge for multimodal large language models (MLLMs) is moving beyond visual recognition to constraint-aware cross-modal reasoning. This involves combining visual cues with information from other modalities to understand elements' relationships under domain-specific rules. This challenge is acutely evident in industrial design-to-code (D2C), which converts user interface (UI) designs into code and requires MLLMs to connect design images with disorganized layer metadata, infer component and layout implementation requirements, and realize them in code under target-library constraints. However, these capabilities remain insufficiently evaluated in realistic industrial settings. To fill this gap, we present TaoD2C-Bench, a benchmark for evaluating MLLMs' ability to generate UI code that satisfies implementation requirements in industrial applications. The TaoD2C dataset consists of 2,861 production designs from 17 commercial platforms with 97,652 expert annotations across four categories: Component, Group, Alignment, and Position. These annotations distinguish required constraints from permitted implementation choices. TaoD2C-Bench defines three tasks: end-to-end UI code generation, requirement inference, and requirement realization. Evaluating eight MLLMs reveals substantial gaps in generating UI code that satisfies implementation requirements, alongside distinct performance profiles in inference and realization. We further show that MLLMs' visual reconstruction ability does not necessarily imply an ability to generate code that meets these requirements. We release TaoD2C to support research on industrial UI code generation.
Create a lesson
Related papers
When Sub-Agents Work in Parallel: The Promises and Pitfalls of Dynamic Concurrency in Long-Horizon Coding Tasks
Han Li, HanHaoNing Li, Ziqian Jiang et al.
Using Small Language Models to Reverse-Engineer Machine Learning Pipelines Structures
Nicolas Lacroix, Frederic Precioso, Mireille Blay-Fornarino et al.
QuSema: Detecting Silent Bugs in Quantum Libraries via Quantum-knowledge-enhanced Agents
Yujin Song, Kaining Zhang, Qixin Zhang et al.
TestGRAD: Evolving Test Suites via Failure Pattern Momentum for SWE-Agent Ensemble
Pengfei He, Jiayuan Zhou, Shaowei Wang et al.
Why Software Engineering Is Indispensable in the Age of Coding Agents
Alfonso Fuggetta
AdaT2: Adaptive Test Transformations for Black-Box Boundary Testing of Conversational Agents
Liting Lin, Boxi Yu, Qinghua Xu et al.