SMM Transformer: Leveraging Spiking Neural Networks for Multimodal Tasks
Xiubo Liang, Jinxing Han, Yuke Li, Haoqi Zhu, Yu Zhao, Hongzhi Wang
Abstract
Spiking Neural Networks (SNNs) enable event-driven computation with sparse activations, but building multimodal Transformers on SNNs is hindered by unstable training in deep spiking stacks and the mismatch between dense softmax attention and spike-based communication. We propose SMM Transformer, an SNN-based multimodal Transformer framework that combines (i)PLMP, a Parallel LIF with Multistage Learnable Parameters neuron and a tailored P-STBP algorithm for stable deep SNN training, (ii) SMSA, an attention-inspired spike-driven token-mixing module that replaces dense pairwise softmax attention with channel-wise spike co-activation and self-compensation, and (iii)SMoE, a spiking mixture-of-experts module for modality-aware fusion. Across visual and multimodal benchmarks, SMM Transformer achieves competitive accuracy compared to ANN baselines. Under a standard MAC/AC arithmetic model, SMSA reduces the estimated operator-level compute energy of the attention module by up to 97%, while whole-model profiling shows more moderate but consistent efficiency gains.
Create a lesson
Related papers
A Metaheuristic Optimization Framework for Discrete Optimization under Strict Time Limits
Umut Çalıkyılmaz, Nitin Nayak, Sven Groppe
Benchmarking Tabular Foundation Models as Surrogates in Expensive Evolutionary Optimization
Lu Han, Jin Wang, Yuchen Li et al.
A Spatiotemporal Extension of the Neuromorphic DBSCAN Implementation
Charles P. Rizzo, James S. Plank
Machine Zygote: Causal Biparental Heredity Before Learning in a Germline--Soma Artificial Agent
Lyes Saad Saoud
Bio-Inspired Palette Evolution in Indirectly Encoded Substrates: Timescale Compatibility Shapes Activation Function Discovery
Romain Claret, Michael O'Neill, Paul Cotofrei et al.
LLMDE: A Large Language Model-Driven Differential Evolution Algorithm for Portfolio Optimization
Rong Chai, Vaclav Snasel, Xiaopeng Wang et al.