Underwater Visual Target Tracking with Target-Specific Depth Estimation and Adaptive Model-Fusion Predictive Control
Yuheng Zhou, Haiyang Cheng, Yanqi Feng, Pangkit Fong, Mei Xuan Lee, Marcus Gee, Chongrong Fang, Jianping He
Abstract
Vision-based underwater target tracking is challenged by unreliable depth measurements and unknown target motion. This paper proposes a stereo visual-servoing framework for an autonomous underwater vehicle (AUV). For perception, the framework derives a stable 3D relative state from stereo images through target-specific depth extraction and Kalman filtering. It constructs a target-depth mask from color, disparity, and temporal cues to select reliable target pixels, and then filters the resulting depth measurement and detected image center separately. For control, the framework decouples yaw regulation from translational control, avoiding computationally expensive coupled multi-DOF optimization and enabling real-time translational MPC. The translational controller employs adaptive model-fusion predictive control, combining constant-velocity and zero-velocity target models to accommodate different target-motion patterns. It updates the model weights using historical prediction errors and computes translational commands subject to actuation, following-distance, and field-of-view constraints. Through simulations and real-world experiments, we validate the effectiveness of the proposed framework and show it has better performance than existing frameworks.
Create a lesson
Related papers
GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies
Xin Chen, Sen Chen, Yujuan Ding et al.
Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control
Hanchu Zhou, Brendan Lynch, Raman Goyal et al.
OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher
Damiano Da Col, Maximilian Igl, Peter Karkus et al.
MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement Learning in Unstructured Autonomous Driving
Thomas Steinecker, Denis Trescher, Alexander Bienemann et al.
MoWAM: Explicit Future Motion Prediction for Efficient World Action Models
Jiayu Wang, Bin Zhu, Yue Yu et al.
HOPHY: A Hierarchical Hypergraph Representation for Off-Road Path and Mission Planning
Pranay Meshram, Charuvahan Adhivarahan, Prithvi Poddar et al.