InfraPatch: Cross-Task Targeted Grayscale Patch Attacks on Infrared-Adapted Vision-Language Models
Chengyin Hu, Dingyi Lu, Jiaju Han, Xiang Chen, Weiwen Shi, Jiahuan Long, Yiwei Wei, Jiujiang Guo
Abstract
Infrared vision-language models (IR-VLMs) have emerged as a promising paradigm for multimodal perception under low-visibility conditions, yet their robustness to targeted adversarial attacks remains poorly understood. Existing adversarial patch methods mainly study RGB-based models or a single downstream task and do not characterize whether localized perturbations can induce an intended semantic target in IR-VLMs. We propose InfraPatch, a white-box, per-instance framework for targeted digital grayscale patch attacks against IR-VLMs. InfraPatch optimizes a compact single-channel patch within an approximately 5% local-area budget, combines proxy-guided placement with task-adaptive semantic objectives, and induces target behaviors in image classification, image captioning, and binary visual question answering. We evaluate ten infrared-adapted model variants on 300 synthetic infrared-style images generated by applying DiffV2IR to a fixed 30-category COCO subset, using clean-conditioned targeted success criteria. InfraPatch achieves targeted attack success rates from 86.00% to 100% across the ten variants. On CLIP and BLIP-2, proxy location search improves success by 6.67 and 10.33 percentage points over optimized random placement, respectively; LLaVA-1.5 remains saturated near 100% under both settings. Patch-area and objective ablations further expose substantial differences in vulnerability across architectures and task formats. These results show that small grayscale patches can inject chosen target semantics across IR-VLM families under a controlled digital threat model, motivating stronger robustness evaluation for infrared multimodal systems.
Create a lesson
Related papers
SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models
Junchao Huang, Guian Fang, Shengju Qian et al.
Thinking in Pictures: A Systematic Benchmark for Reasoning-driven Image Generation
Yutong Liu, Nan Huang, Xu Cao et al.
PlantC2USeg: Cross-Scale Consistent Pre-Training for Few-Shot Unified Plant Point Cloud Segmentation
Yu Tian, Xintong Jiang, Jan Franklin Adamowski et al.
MuyBridge: Mobile Human Center-of-Mass Estimation from Monocular Video via Sparse Fusion
Aidan Bradshaw, Marco Giordano, David Rode et al.
RoGe: Novel View Synthesis via End-to-End Implicit Reconstruction and Generation
Xiaolei Lang, Ze Kang, Zehao Huang et al.
Efficient All-in-One Weather Restoration using Spectral Harmonization
Paula Garrido-Mellado, Daniel Feijoo, Yuning Cui et al.