A new attack produced more faithful images from classifiers when it had full access to their internal settings.
The study presents Diff-MI, a two-stage attack designed for situations in which an attacker has full access to an AI classifier. It first trains a conditional image generator using public images and labels inferred from the classifier, then repeatedly reconstructs images using both the generator and information from the classifier.
Across datasets and model types, the method improved the visual fidelity of reconstructed images compared with existing approaches. The authors report an average 20% reduction in FID, a measure in which lower scores generally indicate that generated images are closer to real ones, while keeping competitive attack accuracy.
More faithful reconstructions
Diff-MI produced more faithful reconstructions than existing model-inversion attacks in experiments across several datasets and classifier models. The authors report an average 20% decrease in FID compared with state-of-the-art methods, while attack accuracy remained competitive. On the CUB-200-2011 bird dataset, the method outperformed the tested baselines on all reported measures and reduced FID by more than 50% compared with them. On the ChestX-Ray medical dataset, it achieved nearly the best results across the reported measures and reduced FID by an average of 45%. The reconstructed bird images captured details such as beak shape, wing coloration and tail length, according to the study’s visual comparisons.