Importance Scoring of Transformer Attention Heads in Learning Tabular Data
Ahmad Jad Allah, Kazi F. Akhter, Md. Kamrozzaman Bhuiyan, Manar D. Samad
Abstract
Computationally demanding and opaque deep learning models can be better understood and optimized by analyzing how they transform data. While deep transformers have been widely studied in computer vision and natural language processing, their application in tabular data remains relatively underexplored. This paper presents one of the first applications of an importance-scoring metric to interpret multi-head transformer models in learning from tabular data. Experiments conducted on 40 diverse tabular datasets demonstrate robustness to head drops based on the proposed head importance score. In 72.5\% of experimental examples, the model remains most resilient to performance drops when heads with the lowest importance scores are gradually removed. In contrast, removing the most important attention head first results in the greatest reduction in classification performance. A closer look at individual head importance scores across six attention layers reveals that important heads are scattered across layers, with no consistent layer-specific trends. In contrast to the image and language domains, the importance of individual attention heads varies considerably across tabular datasets with different schemas and feature spaces. The proposed importance score can improve efficiency and redundancy within transformer architectures. We make the source code for measuring the importance of individual attention heads publicly available.
Create a lesson
Related papers
Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
Yunpeng Ba, Zhi Zheng, Yue Xie et al.
Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting
Xinwei Qiang, Xiang Fang, Chang Chen et al.
QuantumBoostNet: A Hybrid Classical-Quantum Architecture for Enhanced Accuracy in Cardiac Ultrasound View Identification
Mihai Udrescu-Milosav, Stefan-Alexandru Jura, Mihai Udrescu et al.
MM-Spectrum: Multimodal Multi-spectral Molecular Structural Elucidation with a Stable MoE Framework
Hai-tao Yu, Nan Min, Zheng Fang et al.
Making Latent Evolution Explicit: Operator-Structured Transitions for World Action Models
Xiaoxiao Lu, Yunlong Dong, Jiahao Shi et al.
Circuit Condensation: Post-Training that Concentrates a Behavior's Causal Circuit
Sai Adith Senthil Kumar