Skip to content

On the performance of pre-trained vision transformers for supernova spectral classification using different spectral representations

J. Serrano Bell, P. Gálvez Molina, V. Contreras Rojas, W. Fox Fortino, M. Rojas, P. Protopapas, F. Bianco

astro-ph.IMarXiv:2608.21520

Abstract

The spectroscopic classification of supernovae is a key component of time-domain astronomy and plays an important role in the identification of Type Ia events. The increasing volume and diversity of spectroscopic data motivate the development of automated classification approaches. In this work, we explore the use of Transformer-based vision models for supernova spectral classification, focusing on how different visual representations of spectra and fine-tuning strategies affect classification performance when image-based architectures are applied to intrinsically one-dimensional data. We consider a dataset of 4,011 supernova spectra, augmented and rebalanced into three astrophysically motivated classes: normal Type Ia, other Type Ia subtypes, and core-collapse supernovae. Spectra are encoded as line-plot images using direct flux-wavelength visualizations as well as alternative flux-difference versus wavelength-difference heatmap representations designed to emphasize differential spectral structure. We evaluate several pre-trained Transformer architectures, including Vision Transformers (ViT), Swin Transformers, and a DINOv3-pretrained ViT, and examine the effects of fine-tuning depth, visual representation, and hyperparameter choices. We find that a plain flux-wavelength line plot outperforms the purpose-built maps on average, though log-scale maps remain competitive for specific architecture and fine-tuning combinations. On the held-out test set, our best model (vitb-p16) reaches a macro-F1 score of 86.3%, with per-class F1 scores of 94% for normal Ia, 77% for other Ia subtypes, and 88% for core-collapse supernovae. The dominant misclassification arises from confusion between Ia-91T and normal Ia spectra. These results demonstrate that Transformer-based vision models can provide competitive performance for single-spectrum supernova spectral classification.

Create a lesson