ScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generation
Jiawei Zhang, Hongsong Wang, Pan Zhou
Abstract
Text-driven 3D indoor scene generation has advanced from dataset-bound layout modeling to open-vocabulary synthesis with large language and vision-language models. Yet existing methods remain limited: one-pass generators often yield geometrically invalid layouts, heavy post-hoc optimization is costly and unstable, and prompt-only planners lack reusable layout priors for functional grouping and object relations. We propose ScenePilot, a retrieval-augmented Grow-and-Repair framework that formulates scene generation as prior-guided incremental growth with learned rectification. Given a prompt, the Hierarchical Retrieval-Augmented Planning (HRAP) module retrieves room-, group-, and anchor-level layout priors to support functional group planning. A text-driven base generator then inserts object groups sequentially, while the Reinforcement Multimodal Repair (RMR) module performs lightweight local correction after each insertion and a final global repair after completion. To train this policy, we construct SceneReverse-17k, a repair-trajectory dataset built by perturbing high-quality 3D scenes in position, rotation, and scale, then using inverse operations as executable rectification targets. The policy predicts structured move--rotate--scale actions from rendered views, scene state, retrieved priors, and edit history. By combining HRAP with RMR, ScenePilot offers an efficient alternative to one-shot generation and heavy full-scene optimization, improving physical plausibility, functional coherence, and controllability while preserving diversity.
Create a lesson
Related papers
PRISM: Predictive Recomposition via Semantic Latent Decomposition for View-invariant Video Representation Learning
Youngchae Chee, Hosu Lee, Sungjune Park et al.
MCSeg: Pre-training and Fine-tuning Volumetric Pyramid Transformer for Multi-modal Cardiac Image Segmentation
Zhiyu Ye, Hairong Zheng, Tong Zhang
Seeing the Unseen: Camouflaged Object Detection Beyond the Visible Spectrum
Avi Gupta, Trasha Gupta
Proximity3D: Shape from Capacitive Proximity on Sensing Manifold
Hao Chen, Chenming Wu, Chun Ping Lam et al.
CapFrame: Text-Instructed Viewpoint Grounding in 3D Gaussian Scenes via Geometric Pseudo Labels
Jirong Li, Satoshi Ikehata, Shuhei Kurita et al.
Knowing Beyond the Known: Reinforced Knowledge Specification for Multi-Label Class-Incremental Learning
Aoting Zhang, Dongbao Yang, Chang Liu et al.