Enhancing instance discrimination for self-supervised adversarial training
Abstract
A learning machine is prone to be attacked by adversarial examples in presence of imperceptible perturbations. Adversarial training (AT) is developed to build a defense model against the adversarial attacks. Most of the previous works were conducted in supervised setting. To release the demand of data labeling, the self-supervised learning is required to carry out the unsupervised AT. This study presents a contrastive view in AT where the unsupervised method is strengthened by enhancing the instance discrimination based on contrastive learning (CL). The unsupervised AT is newly implemented by connecting to a self-supervised method. Such a connection is utilized to consolidate the unsupervised AT in accordance with a classification perspective in a contrastive objective. In particular, the issues of the inconsistent class definition in instance discrimination for CL and the inconsistent distribution bias in the data augmentation methods in previous CL and AT are tackled. A novel self-supervised AT is developed to train feature representations for various types of data. Therefore, a self-supervised AT with contrastive view is performed to calculate the perturbed image embeddings to train a robust classifier. The contrastive objective is also re-formulated to train a self-supervised sentence encoder as the defense model. The experiments on image and text representations show that this method improves the previous unsupervised methods, and surpasses some other supervised method.
// Source
Authors: Jen‐Tzung Chien, Yuan-An Chen
Institutions: National Yang Ming Chiao Tung University