Multi-label Classification of Ground-based Auroral Emission Spectra Using 1D Vision Transformers

Jul 6, 2026·
Matthieu Le Lain
Matthieu Le Lain
,
Gaël Cessateur
,
Sébastien Lefèvre
· 1 min read
Abstract
We examine whether Vision Transformer attention can serve as a built-in interpretability mechanism on ordered 1D scientific signals, where physical ground truth enables direct validation. We investigate this question through the multi-label classification of ground-based auroral emission spectra, a novel task with no prior baseline. We compare a 1D Vision Transformer (ViT-1D) with an MLP baseline on 719 expert-annotated spectra spanning five emission classes, and find that both architectures achieve strong performance (MLP macro AP = 81.1%; ViT-1D = 77.8%). Beyond classification, we show that the ViT-1D’s attention maps recover known emission-line locations without spectroscopic priors, demonstrating that Transformer attention provides built-in physical interpretability on ordered 1D sequences which is a property unavailable from non-attentive architectures.
Type
Publication
Reconnaissance des Formes, Image, Apprentissage et Perception (RFIAP 2026)
publications

Paper presented at RFIAP 2026 (Reconnaissance des Formes, Image, Apprentissage et Perception), July 6-8, 2026, Montpellier, France. A joint work with the Royal Belgian Institute for Space Aeronomy, based on auroral spectra acquired by the ASIS spectrograph in Skibotn, Norway.

An interactive demo runs the published 15-model ViT-1D ensemble on a spectrum you upload, returning per-class probabilities and the attention map over your own spectrum.

This work has since been extended to the full ASIS archive of 328,281 spectra, presented at the 22nd International EISCAT Symposium in Kiruna, Sweden.

Matthieu Le Lain
Authors
AI for astronomy & astrophysics
PhD student at IRISA, Université Bretagne Sud (expected 2026), and lecturer in computer science, working on foundation models for astronomy and astrophysics, with a broader interest in deep learning applied to scientific data. Also contributing to the UniverseTBD collaboration on vision-language models for astronomy.