Toward Automated Detection of Data Poisoning and Backdoor Attacks in Healthcare Imaging AI: A Proposed Detection Methodology and Evaluation Framework
Abstract
Machine learning models are increasingly deployed in healthcare imaging pipelines for diagnosticsupport, yet the security of these models against training-time attacks — particularlydata poisoning and backdoor insertion — remains substantially under-tested relative to theirgrowing deployment. Recent healthcare-sector guidance (Health Sector Coordinating CouncilCybersecurity Working Group, 2025) and federal policy (Executive Order 14409, 2026) have explicitlynamed data poisoning and model manipulation as threats requiring dedicated defenses,yet no widely adopted, openly available detection tooling validated specifically on healthcareimaging benchmarks currently exists. This paper proposes a detection methodology combiningspectral signature analysis (Tran et al., 2018) and activation clustering (Chen et al., 2018) —two established backdoor-detection techniques not yet systematically benchmarked together onpublic healthcare imaging datasets — and specifies an evaluation protocol against syntheticallypoisoned variants of public medical imaging benchmarks alongside a non-healthcare criticalinfrastructure-relevant dataset, to test cross-sector generalizability. We frame the resultingevidence output against the NIST AI Risk Management Framework’s Measure function and theMITRE ATLAS adversary technique taxonomy, so that results are directly usable by healthcaresecurity teams rather than requiring translation from research notation. This paper presents themethodology and evaluation plan; empirical results will be reported in a follow-up publication.
// Source
Authors: Suresh Tamang
Institutions: University of the Cumberlands