Improving Generative Deep Learning-Based SAR-to-Optical Image Translation through Spatiotemporal Feature Incorporation: A Case Study in Cropland
Abstract
Synthetic aperture radar (SAR)-to-optical (S2O) image translation is an effective approach for reconstructing cloud-contaminated regions in optical imagery using corresponding SAR observations.However, its ability to reconstruct multispectral optical information is inherently limited because high spatial resolution SAR imagery is predominantly acquired in single-polarization modes that provide only limited information for reliable spectral reconstruction.To address this limitation, this study applied generative deep learning-based S2O image translation over cropland using spatial features derived from wavelet decomposition and temporal features based on day-of-year (DOY) encoding.Four input configurations were analyzed: (1) HH single-polarization SAR imagery alone, (2) HH SAR imagery combined with spatial features derived from the discrete wavelet transform, (3) HH SAR imagery combined with DOY-based temporal features, and (4) HH SAR imagery combined with both spatial and temporal features.Conditional generative adversarial networks (cGAN) and conditional denoising diffusion probabilistic models (cDDPM) were applied to evaluate the consistency of the effects across different generative mechanisms.Experiments were conducted on the cropland in Gimje using COSMO-SkyMed SAR and PlanetScope optical imagery.The results demonstrated that incorporating spatiotemporal features consistently improved the prediction accuracy over the SAR-only case for both models, with average relative improvements of 14.66% and 15.88% for cGAN and cDDPM, respectively.In particular, the contribution of temporal features was more pronounced than that of spatial features because the DOY encoding effectively captured the pronounced phenological changes in rice growth stages within the cropland.These findings suggest that the input feature configuration plays a critical role in generative deep learning-based S2O image translation performance, and multispectral optical imagery can be reconstructed effectively by incorporating additional spatiotemporal features, even when only single-polarization SAR data are available.
// Source
Authors: Soyeon Park, M. Park, Eui-Ho Hwang, No-Wook Park
Institutions: Inha University