The study introduces the Friend-Safe Adversarial Attack, which creates altered inputs designed to cause errors in selected models while keeping other models’ predictions largely unchanged. The researchers tested the approach on MNIST, CIFAR-10 and SVHN, as well as additional personalization methods, model pairings and image datasets.

With exact access to the friend models, the attack reached at least an 88% success rate against target models and maintained at least 81% accuracy on friend models. When the attacker instead used information available to a passive aggregation server, friend-model predictions were preserved by 92–96% on MNIST, 70–74% on SVHN and 37–43% on CIFAR-10.