FunArt: Decoding Functional Structure and Articulation from Generative 3D Latents
Dennis Rotondi, Abdelrhman Werby, Kai O. Arras
Abstract
To operate effectively in human environments, robots must identify articulated objects, segment their movable and interactive parts, and estimate their kinematic models. Existing articulated scene representations typically recover kinematics from observed interactions, while methods operating on static scans often decouple articulation from functional interactive elements. We present FunArt, a framework that constructs articulation-aware functional 3D scene graphs from posed RGB-D observations captured in a single static configuration. FunArt reconstructs object instances, converts their fused geometry directly into the O-Voxel representation of TRELLIS.2, and exploits its frozen, sparse-compression VAE as a structural prior. A lightweight query-based decoder combines compact object-level latents with dense, surface-aligned features to jointly segment movable parts and functional interactive elements while estimating motion type, axis, origin, and range. On the Articulate3D dataset, FunArt achieves state-of-the-art performance across movable-part segmentation, articulation estimation, and functional-element segmentation, both with and without ground-truth object input. In the end-to-end setting, it outperforms the strongest baselines by 1.5 AP50 points for movable parts, 2.8 AP50 points under joint origin-and-axis constraints, and 6.7 AP50 points for functional elements. These results demonstrate that generative 3D latents encode actionable structural cues that can initialize robotic perception and planning before physical interaction.
Create a lesson
Related papers
FlowSGS: Improving Flow Matching Priors for Inverse Imaging with Stochastic Interpolants
Tianao Li, Xinhui Qian, Emma Alexander
Should This Case Be Adapted? Prediction Fragmentation Controls Test-Time Adaptation
Lili Wang, Jing Li, Xiaowen Sun et al.
Earth Surface Immune System for Rapid Monitoring of Unknown Anomalies
Jingtao Li, Qian Zhu, Xinyu Wang et al.
PROVIA: Procedure State Tracking for Online Mistake Detection in Egocentric Videos
Di Wen, Kailun Yang, Jimmy Weissert et al.
Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network
Yulong Chen, Ziqian Zhang, Haoyu Zhang et al.
PhGS: Post-Hoc Pruning and Refinement of Single-View Feed-Forward 3D Gaussian Reconstructions
Rinto Yagawa, Han Cheng, Dieter Schmalstieg et al.