AI & Computingarticle2026-08-14

A diffusion-based multi-style image generation framework with adaptive weighted style fusion

Open access0 citations

Abstract

Diffusion-based multi-style image generation aims to preserve the semantic structure of a content image while integrating visual characteristics from multiple style references. Existing style transfer and diffusion-based methods often rely on single-style guidance, fixed prompts, or simple feature fusion, which may cause style conflict, uneven style dominance, and weakened content preservation. To address these problems, this study proposes a diffusion-based multi-style image generation framework with adaptive weighted style fusion. A pretrained VGG encoder is first used to extract multi-level content and style features. Then, an adaptive weighting mechanism calculates the contribution of each style reference according to style complexity and style-content similarity. A cross-attention-guided fusion module is introduced to enhance content-aware style interaction, while an adaptive modulation module is embedded into the denoising U-Net to dynamically control style influence during reverse denoising. Experiments on 1000 content images and 1000 style images show that the proposed method achieves PSNR of 32.4 dB, SSIM of 0.91, FID of 24.3, LPIPS of 0.176, CLIP-I of 0.846, and style consistency score of 0.834. Visual comparison, ablation study, and user evaluation confirm that the framework improves content preservation, perceptual quality, and multi-style coordination.

// Source

View paper (DOI)Open access versionOpenAlexDiscover Artificial IntelligencePublished 2026-08-14

Authors: Haoran Gong

Institutions: Nanjing University of Information Science and Technology