CapFrame: Text-Instructed Viewpoint Grounding in 3D Gaussian Scenes via Geometric Pseudo Labels
Jirong Li, Satoshi Ikehata, Shuhei Kurita, Ikuro Sato
Abstract
3D Gaussian Splatting (3DGS) enables photorealistic real-time novel view synthesis, yet placing a virtual camera to capture a desired frame remains largely manual. Existing language-guided approaches in 3D scenes mainly focus on object-centric grounding, determining what to observe but rarely controlling how it should appear in a single frame, such as subject orientation or frame layout. To address this limitation, we introduce a new task, Text-Instructed Viewpoint Grounding (TIVG), which aims to identify a 6-DoF camera pose in a 3D Gaussian scene whose rendered frame aligns with a text instruction. To solve this task, we propose CapFrame, a partially differentiable framework that converts language into geometric pseudo labels for camera pose optimization. CapFrame follows a Retrieve-Translate-Refine pipeline: it retrieves relevant views and ranks them through a Question-Evaluation process with MLLMs, translates the instruction into orientation and layout pseudo labels, and refines the camera pose via differentiable optimization with layout and orientation losses in 3DGS. Experiments on 38 real-world scenes with 135 instructions indicate that CapFrame produces viewpoints better aligned with texts than heuristic viewpoint search and adapted trajectory generation baselines, validated by VLM metrics, MLLM judges, and user studies. Code is available at: https://github.com/jirongli/CapFrame
Create a lesson
Related papers
PRISM: Predictive Recomposition via Semantic Latent Decomposition for View-invariant Video Representation Learning
Youngchae Chee, Hosu Lee, Sungjune Park et al.
MCSeg: Pre-training and Fine-tuning Volumetric Pyramid Transformer for Multi-modal Cardiac Image Segmentation
Zhiyu Ye, Hairong Zheng, Tong Zhang
Seeing the Unseen: Camouflaged Object Detection Beyond the Visible Spectrum
Avi Gupta, Trasha Gupta
Proximity3D: Shape from Capacitive Proximity on Sensing Manifold
Hao Chen, Chenming Wu, Chun Ping Lam et al.
Knowing Beyond the Known: Reinforced Knowledge Specification for Multi-Label Class-Incremental Learning
Aoting Zhang, Dongbao Yang, Chang Liu et al.
ScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generation
Jiawei Zhang, Hongsong Wang, Pan Zhou