The method creates substitute image-and-text examples by blending features and adding calibrated noise in the prompt rather than changing the model's parameters. This is intended to hide fine details that could identify sensitive examples while retaining broader patterns needed for classification, under a formal differential privacy guarantee.

Tests used a proprietary smart-grid inspection dataset and a subset of the MIMIC-CXR chest-radiograph dataset. In five-shot experiments, the method outperformed the comparison systems at the privacy settings reported, while using the frozen LLaVA-1.5-7B multimodal model rather than training an external generator.