Text2Villa: Hierarchical Generation of 3D Indoor Environments with Physics-Aware Analysis-by-Synthesis
Xiang Tang, Ruotong Li, Xiaopeng Fan
Abstract
Generating 3D indoor scenes from natural language holds tremendous potential, yet existing methods predominantly fail to generate multi-room structures with vertical connectivity and arbitrary polygonal boundaries. Furthermore, they lack a deep grounding in continuous 3D physical laws, leading to severe geometric penetrations and floating artifacts. In this work, we propose Text2Villa, a novel hierarchical generative framework. At the macro level, we construct a multi-story dataset to fine-tune an autoregressive layout generator, ensuring the direct parsing of text into 3D building foundations featuring polygonal boundaries and multi-story connectivity. To enforce physical laws during micro-level asset arrangement, we introduce the Affordance-driven Physical-Semantic Scene Graph (A-PSSG) to explicitly abstract physical affordances (such as support surfaces and containment cavities) into node attributes, establishing strict geometric and semantic edge constraints. Guided by the A-PSSG, we formulate scene instantiation as a constrained closed-loop optimization problem following the analysis-by-synthesis paradigm. By integrating an underlying geometric collision detection engine with the high-level semantic reasoning of multimodal large language models (MLLMs), our heuristic solver dynamically executes physics-aware actions under the observation-evaluation-modification mechanism to effectively resolve mesh collisions, floating artifacts, and fine-grained cavity containment failures. Extensive experiments demonstrate that Text2Villa outperforms previous methods across various metrics, robustly generating high-fidelity and physically plausible villa-level 3D environments from text, thereby providing a reliable and interactive 3D content foundation for downstream applications.
Create a lesson
Related papers
RenderFormer++: Scalable and Physics-Informed Feed-Forward Neural Rendering
Huangsheng Du, Haoran Zhu, Youcheng Cai et al.
ESVR: 3D Ellipsoid-based Sparse Volume Rendering via Structure-aware Primitive Learning and Per-primitive Ray Sampling
Suemin Jeon, Youjin Kim, Jungwoo Park et al.
GenGA: Editable and Data-Grounded Graphical Abstract Generation for Academic Papers
Takuro Kawada, Shunsuke Kitada, Hitoshi Iyatomi
Fourier-Latent Diffusion for Constrained Generation of Triply Periodic Minimal Surfaces
Shu Yan, Bohan Wang
URHead: A Unified UV-Space Representation for Joint Mesh-3DGS Optimization in Head Avatars
Seonghak Lee, Junhee Cho, Jisoo Park et al.
Style-Aware Gloss Control for Generative Non-Photorealistic Rendering
Santiago Jimenez-Navarro, Belen Masia, Ana Serrano