Deep generative models for enhanced vitreous OCT imaging
Abstract
Abstract The aim of the work is to evaluate deep learning (DL) models for enhancing vitreous optical coherence tomography (OCT) image quality and potentially reducing acquisition time. Conditional Denoising Diffusion Probabilistic Models (cDDPMs), Brownian Bridge Diffusion Models (BBDMs), U-Net, Pix2Pix, and Vector-Quantised Generative Adversarial Network (VQ-GAN) were used to generate high-quality spectral-domain (SD) vitreous OCT images. Inputs were SD ART10 images, and outputs were compared to pseudoART100 images obtained by averaging ten ART10 images per eye location. Model performance was assessed using image quality metrics and Visual Turing Tests, where ophthalmologists ranked generated images and evaluated anatomical fidelity. The best model’s performance was further tested within the manually segmented vitreous on newly acquired data. U-Net achieved the highest Peak Signal-to-Noise Ratio (PSNR: 30.230) and Structural Similarity Index Measure (SSIM: 0.820), followed by cDDPM. For Learned Perceptual Image Patch Similarity (LPIPS), Pix2Pix (0.697) and cDDPM (0.753) performed best. In the first Visual Turing Test, cDDPM ranked highest (3.07); in the second (best-ranked model only), performed as exploratory analysis, cDDPM achieved a 32.9% fool rate and 85.7% anatomical preservation. On newly acquired data, cDDPM generated vitreous regions significantly more similar in SSIM to the ART100 reference than true ART1 or ART10 B-scans and achieved similar scores on whole images. Results reveal discrepancies between quantitative metrics and clinical evaluation, highlighting the need for combined assessment. Only SSIM was significantly positively correlated with clinical judgement. Overall, cDDPM showed potential for generating clinically meaningful vitreous OCT images. Dataset and code are made publicly available.
// Source
Authors: Simone Sarrocco, Philippe C. Cattin, Peter M. Maloca, Paul Friedrich, Philippe Valmaggia