The Illusion of Generalization: Impact of Ultra-Low Synthetic Contamination on Deep Waveform AI Scratch Training in Acoustic Leak Detection
Abstract
This paper investigates the fundamental problem of false adaptation in deep neural networks applied to non-destructive testing (NDT) of pipelines. Under severe field data shortages, developers often rely on surrogate audio files generated by rapid, heuristic methods (such as LLM-assisted vibe-coding) tailored to the parameters of physical correlation leak detectors. We experimentally demonstrate that injecting as little as 1% of such synthetic data into scratch training of a deep edge waveform classifier creates a robust illusion of model improvement: on internal validation, accuracy increases from 97.06% to 99.67%, while false positives drop ninefold. However, when tested on an independent, out-of-domain field dataset (Zayed et al., WDN Hong Kong, 140 files), the model exhibits a 7.14% drop in Recall. As synthetic contamination increases to 20%, a complete collapse of generalization capability occurs — Recall plummets from 70.00% to 11.43%. This phenomenon is driven by the network's convolutional filters shifting focus to simple mathematical patterns of the generator (shortcut learning) instead of the complex, noisy physical wave envelopes.
// Source
Authors: Aleksandr Ivanaiskii, Evgeny Ivanaiskii, Shipilov Shipilov
Institutions: Heilongjiang University of Technology