AI & Computingarticle2026-09-02

AI4TEN: Fine-Tuning Pre-Trained Audio Transformers (BEATs, AST) for Cross-Dataset Acoustic Vehicle Classification with Domain Adaptation

Open access0 citations

Abstract

Acoustic vehicle classification from roadside microphones supports source-specific traffic noise monitoring, but classifiers trained on one dataset generalize poorly to recordings from different environments. This study compares pre-trained audio transformers (BEATs, AST) against a CNN baseline for cross-dataset vehicle classification into five categories (car, truck, motorcycle, bus, background), training on the IDMT-Traffic dataset (Germany) and testing on the MELAUDIS dataset (Australia). Two domain adaptation methods (DANN, ArcFace) are applied to the best-performing transformer. BEATs achieves a balanced F1 of 0.46, a 77% improvement over the CNN baseline (0.26). AST achieves F1 = 0.43, a 65% improvement. Both pre-trained transformers substantially outperform the baseline, with self-supervised pre-training producing slightly more transferable representations than supervised pre-training. ArcFace metric learning achieves the highest cross-dataset F1 (0.50), modestly outperforming standard fine-tuning. DANN degrades cross-dataset performance but achieves the highest microphone robustness score (F1 = 0.80). Classification of the underrepresented bus class (53 training samples) improves from F1 = 0.38 to 0.54 with ArcFace. Microphone robustness is strong across all models (BEATs F1 = 0.75), confirming that recording environment dominates the domain shift over hardware variation. The choice of domain adaptation strategy should be guided by the expected type of domain shift.

// Source

View paper (DOI)Open access versionOpenAlexAcousticsPublished 2026-09-02

Authors: Jibran Khan

Institutions: Aarhus University, University of Eastern Finland