Monocular 3D Pose Recovery via Nonconvex Sparsity with Theoretical Analysis

Abstract

For recovering 3D object poses from 2D images, a prevalent method is to pre-train an over-complete dictionary D=\Bi\iD of 3D basis poses. During testing, the detected 2D pose Y is matched to dictionary by Y ≈ Σi Mi Bi where \Mi\iD=\ci Ri\, by estimating the rotation Ri, projection and sparse combination coefficients c ∈ R+D. In this paper, we propose non-convex regularization H(c) to learn coefficients c, including novel leaky capped 1-norm regularization (LCNR), align* H(c)=α Σi (|ci|,τ)+ β Σi (| ci|,τ), align* where 0≤ β ≤ α and 0<τ is a certain threshold, so the invalid components smaller than τ are composed with larger regularization and other valid components with smaller regularization. We propose a multi-stage optimizer with convex relaxation and ADMM. We prove that the estimation error L(l) decays w.r.t. the stages l, align* Pr( L(l) < l-1 L(0) + δ ) ≥ 1- ε, align* where 0< <1, 0<δ, 0<ε 1. Experiments on large 3D human datasets like H36M are conducted to support our improvement upon previous approaches. To the best of our knowledge, this is the first theoretical analysis in this line of research, to understand how the recovery error is affected by fundamental factors, e.g. dictionary size, observation noises, optimization times. We characterize the trade-off between speed and accuracy towards real-time inference in applications.

0

Turn this paper into a lesson

ArcXiv compiles a structured reading guide from this paper's metadata: plain-English importance, contributions, prerequisite concepts, which sections to read first, flashcards, and a quiz. Grounded in the abstract, never invented.

Discussion (0)

Sign in to join the discussion.

Loading comments…