Multi-label Classification of Ground-based Auroral Emission Spectra Using 1D Vision Transformers

Jul 6, 2026·
Matthieu Le Lain
Matthieu Le Lain
,
Gaël Cessateur
,
Sébastien Lefèvre
· 1 min read
Abstract
We examine whether Vision Transformer attention can serve as a built-in interpretability mechanism on ordered 1D scientific signals, where physical ground truth enables direct validation. We investigate this question through the multi-label classification of ground-based auroral emission spectra, a novel task with no prior baseline. We compare a 1D Vision Transformer (ViT-1D) with an MLP baseline on 719 expert-annotated spectra spanning five emission classes, and find that both architectures achieve strong performance (MLP macro AP = 81.1%; ViT-1D = 77.8%). Beyond classification, we show that the ViT-1D’s attention maps recover known emission-line locations without spectroscopic priors, demonstrating that Transformer attention provides built-in physical interpretability on ordered 1D sequences which is a property unavailable from non-attentive architectures.
Type
Publication
Reconnaissance des Formes, Image, Apprentissage et Perception (RFIAP 2026)
publications

Paper presented at RFIAP 2026 (Reconnaissance des Formes, Image, Apprentissage et Perception), July 6-8, 2026, Montpellier, France. A joint work with the Royal Belgian Institute for Space Aeronomy, based on auroral spectra acquired by the ASIS spectrograph in Skibotn, Norway.

Matthieu Le Lain
Authors
AI for astronomy & astrophysics
PhD student at IRISA, Université Bretagne Sud (expected 2026), and lecturer in computer science, working on foundation models for astronomy and astrophysics, with a broader interest in deep learning applied to scientific data. Also contributing to the UniverseTBD collaboration on vision-language models for astronomy.