AI & Computingpreprint2026-08-14

The Illusion of Generalization: Impact of Ultra-Low Synthetic Contamination on Deep Waveform AI Scratch Training in Acoustic Leak Detection

Open access0 citations

Abstract

This paper investigates the fundamental problem of false adaptation in deep neural networks applied to non-destructive testing (NDT) of pipelines. Under severe field data shortages, developers often rely on surrogate audio files generated by rapid, heuristic methods (such as LLM-assisted vibe-coding) tailored to the parameters of physical correlation leak detectors. We experimentally demonstrate that injecting as little as 1% of such synthetic data into scratch training of a deep edge waveform classifier creates a robust illusion of model improvement: on internal validation, accuracy increases from 97.06% to 99.67%, while false positives drop ninefold. However, when tested on an independent, out-of-domain field dataset (Zayed et al., WDN Hong Kong, 140 files), the model exhibits a 7.14% drop in Recall. As synthetic contamination increases to 20%, a complete collapse of generalization capability occurs — Recall plummets from 70.00% to 11.43%. This phenomenon is driven by the network's convolutional filters shifting focus to simple mathematical patterns of the generator (shortcut learning) instead of the complex, noisy physical wave envelopes.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-14

Authors: Aleksandr Ivanaiskii, Evgeny Ivanaiskii, Shipilov Shipilov

Institutions: Heilongjiang University of Technology