Society & Economicsarticle2026-08-18

Lightweight Multimodal-Guided Diffusion Transformer for Agricultural Product Packaging Visual Concept Generation

0 citations

Abstract

To address the problems of high costs during the conceptual design stage of agricultural product packaging and weak correlation between visual styles and market feedback, this study proposes a Lightweight Multimodal-Guided Diffusion Transformer (LMG-DiT). Driven by the Amazon Product Dataset (APD), the model first adopts a Multimodal Large Language Model (MLLM) to extract design-oriented market instructions from massive commercial data. Subsequently, it relies on the Diffusion Transformer (DiT) architecture to maintain the rational spatial layout of packaging pages. On this basis, two lightweight strategies, namely Latent Consistency Distillation (LCD) and 4-bit Integer quantization (INT4), are integrated to enable efficient operation on terminal devices with limited computing resources. Experimental results demonstrate that the visual concepts of agricultural product packaging generated by LMG-DiT achieve outstanding performance across multiple evaluation metrics, with a Fréchet Inception Distance (FID) of 12.51, an Inception Score (IS) of 16.85, and an average inference time of merely 1.5 seconds per solution. Compared with open-source baselines including Stable Diffusion v1.5, Vanilla DiT-XL/2, and Layout-Transformer, the proposed model presents distinct advantages in visual authenticity, commercial semantic matching, and operational efficiency. Overall, this method provides an intelligent solution that balances efficiency, cost control and market adaptability for the visual branding construction of agricultural products in resource-constrained scenarios.

// Source

View paper (DOI)OpenAlexInternational Journal of Pattern Recognition and Artificial IntelligencePublished 2026-08-18

Authors: Yanan Jiang, Yongxiao Liu

Institutions: Twitter (United States)