Mechanism-Level Evaluation for Vision-Language Models: Controlled Activation-Replacement Diagnosis of Gender Bias
Zhipeng Zhao, Wenxu Wang, Peishun Liu, Ruichun Tang
Abstract
Behavioral benchmarking reveals what biases exist in vision-language models but not which internal components are most sensitive to targeted intervention, precluding principled intervention. We argue for mechanism-level evaluation as a necessary complement, demonstrating causal mediation analysis as a diagnostic instrument for gender bias. We decompose gender-cue effects into controlled indirect effects attributable to specific-layer activations and direct effects through all other pathways, producing layer-by-layer mechanistic signatures. Across six models spanning three architectural families (LLaVA-1.5, LLaVA-NeXT, InstructBLIP at 7B/13B) and two 8B-scale architectures, three findings emerge: language-layer activations exhibit the greatest output sensitivity under controlled intervention, with the direct component often carrying the opposite sign; architectural choices redistribute layer-wise sensitivity to activation replacement; and counterfactual scores diverge from surface-level scores, exposing implicit associations. Systematic ablation validates internal consistency. An intervention experiment finds that the average indirect effect (AIE) and downstream intervention effectiveness are only weakly correlated (Pearson r = 0.33), and the layer with the second-largest AIE produces near-zero bias change---indicating that mechanistic diagnosis captures activation-replacement sensitivity but does not, by itself, identify optimal intervention targets. These results show mechanism-level evaluation captures architecture-specific sensitivity patterns that behavioral benchmarks cannot; pairing both should become standard NLP practice. Code: https://github.com/zhaozhipeng1997/CARD-GenderBias.
Create a lesson
Related papers
MoQSplat: Adaptive Progressive Streaming of 3D Gaussian Splatting via MoQ
Emanuele Artioli, Mohammadreza Ghafari, Md Tariqul Islam et al.
Divide and Conquer: Mixture-of-Bottleneck Experts in Informative Ordinal Space for Video-based Multimodal Sentiment Analysis
Ronghao Lin, Qiaolin He, Zefeng Lu et al.
Multimodal Aspect-Level Sentiment Analysis Based on Gated Noise Filtering and Emotion-Relevance Interaction
Chen Huang, Liangwei Guo, Yamin Li et al.
SemABR: Measuring Video Semantic Fidelity with Multimodal LLMs for Adaptive Bitrate Streaming
Shiqi Xu, Soung Chang Liew, Yuyang Du
Multimodal Emergency Vehicle Classification via Audio-Visual Transformers and Knowledge Distillation
Vijay John, Amar Dabaja
A Low-Latency Interactive System for Real-Time Video Understanding Based on VLMs
Punan Dai, Jun Xu, Bingcong Lu et al.