An Energy-Proportional Multimodal and Context-Aware Vision IoT Node
Julian Moosmann, Philipp Mayer, Luca Benini, Michele Magno
Abstract
While recent advancements in TinyML have significantly reduced the computational complexity of on-device vision pipelines, image acquisition remains a dominant contributor to system-level energy consumption and memory footprint. In vision-enabled IoT platforms, the image sensor consumes energy comparable to the inference engine, thereby offsetting algorithmic efficiency gains. Consequently, current designs face a fundamental trade-off: continuous and always-on sensing incurs prohibitive energy consumption, whereas aggressive duty cycling increases latency and risks missing transient events. This work presents an energy-proportional, context-aware vision IoT node that addresses this challenge through a heterogeneous multimodal dual-camera architecture. Detection and recognition are decoupled by combining an event-based imager operating asynchronously in an energy-efficient always-on wake-on-motion mode together with an RGB imager. Deployed on a low-power microcontroller, a novel TinyissimoYOLOv12 is introduced for efficient and accurate object detection. By activating the high-power image acquisition and processing stages only upon sparse visual triggers, the proposed architecture improves efficiency and latency, eliminating redundant sensing while maintaining continuous monitoring coverage. Experimental results demonstrate an energy consumption of only 222μWh. Upon a motion trigger, the system completes a full sense-to-report cycle-RGB acquisition, object detection across 80 classes, and LoRa telemetry-with a total energy consumption of 28.7mJ. The network achieves up to 32.3% mAP with a model size of 1 million parameters. At a 1% daily activity ratio, the platform achieves a three-month operational lifetime with a 1.85Wh battery, enabling always-on visual monitoring in a place-and-forget scenario through autonomous edge intelligence.
Create a lesson
Related papers
Highly accelerated 3D Cartesian MPnRAGE with implicit neural representation reconstruction
Natascha Niessen, Ana Beatriz Solana, Carolin M. Pirkl et al.
Informed Sinogram Interpolation for Sparse View Reconstruction
Yuejie Liu, Alessandro Lupoli, David Uribe Gallo et al.
Constrained Color Carrier: Characterization-Preserving Conditional Color Rendering in Multi-Illuminant Camera Profiles
Xilai Liang
Quantum-Inspired Trainable and Parameter-Efficient Tensor Networks for Image Inpainting
Shiwen An, Konstantinos Slavakis
StainBridge: Stain-Aware Pairwise Registration of Serial Renal Biopsy Whole-Slide Images Across Structural and Immunohistochemical Stains
Ellen Wei, Bohang Jiang, Yanfan Zhu et al.
Semantic-Aware Neural Video Codec for Error-Resilient Low-Latency Transmission
Matin Mortaheb, Homa Esfahanizadeh, Jinfeng Du et al.