Health & Medicinearticle2026-08-03

Development and validation of CVD risk prediction models based on food quality in an Iranian population using real world data from the Persian cohort of Sabzevar: A comparative study of several machine learning algorithms.

Open access0 citations

Abstract

Project Title: Development and Validation of CVD Risk Prediction Models Based on Food Quality in an Iranian Population Using Real-World Data from the Persian Cohort of Sabzevar: A Comparative Study of Several Machine Learning Algorithms 1. Purpose and Rationale The primary purpose of this research is to develop and internally validate a novel, robust, and accurate cardiovascular disease (CVD) risk prediction model specifically for the Iranian population. Cardiovascular diseases are the leading cause of death globally and in Iran, creating an immense burden on healthcare systems. While numerous CVD risk prediction models exist (e.g., Framingham Risk Score, QRISK, WHO risk charts), they are often developed in Western populations and may not perform optimally in other ethnic groups due to differences in lifestyle, genetics, and environmental factors. This project aims to address this critical gap by leveraging two key innovations: Integration of Dietary Quality: The project is unique in its integration of a comprehensive Food Quality Score (FQS) as a central predictive factor. While traditional models focus on clinical and demographic factors (e.g., age, cholesterol, blood pressure), this study hypothesizes that diet quality provides crucial, independent predictive power. By using the FQS—which evaluates the intake of beneficial and harmful food groups—the model aims to capture the nuanced impact of nutrition on CVD risk. Application of Machine Learning: Instead of relying solely on traditional statistical methods like logistic regression, this study will employ and compare several advanced machine learning (ML) algorithms. This includes Support Vector Machines (SVM), k-Nearest Neighbors (KNN), Multi-Layer Perceptron (MLP), and Random Forest. Machine learning is adept at identifying complex, non-linear relationships within large datasets, potentially leading to a more accurate and personalized risk assessment than conventional models. Ultimately, this research is designed to enhance precision public health strategies in Iran by providing a predictive tool that is both highly accurate and relevant to the local population. The ultimate goal is to facilitate early identification of at-risk individuals, enabling timely and effective preventative interventions to reduce the national burden of CVD. 2. Methodology and Approach This prediction model study will be conducted using baseline data from the Sabzevar Persian Cohort Study (SPCS), which is part of the larger PERSIAN cohort. The SPCS enrolled 4,218 permanent residents of Sabzevar, Iran, aged 35-75 years, providing a rich and representative real-world dataset. The research is structured around two specific aims and will be developed in a stepwise, systematic manner: Aim 1: Model Development: Phase 1: Problem Definition: The primary outcome is the binary prediction of CVD risk. The success of the model will be measured using performance metrics like accuracy, sensitivity, specificity, and the Area Under the Receiver Operating Characteristic Curve (AUC). Phase 2: Variable Identification and Data Preparation: Predictors: A comprehensive set of predictors will be identified through a systematic literature review. Beyond standard clinical and demographic variables, the project will compute the Food Quality Score (FQS) for each participant. The FQS (scored 14-70) is based on 14 food groups, where higher scores indicate a healthier diet. Feature Selection: The "Relief" feature selection algorithm will be used to identify the most significant predictors, ensuring the model is efficient and avoids overfitting. Data Handling: The project outlines clear plans to address common data challenges. A hybrid approach (combining oversampling of the minority class and class weighting) will be used to handle any class imbalance in CVD cases. Multiple imputation will be used to address missing data. Phase 3: Model Selection: A variety of ML models (SVM, KNN, MLP, Random Forest) will be developed and trained on the dataset. The model's performance will be compared to determine the most effective algorithm. Phase 4: Model Performance Reporting: The performance of each model will be meticulously reported and compared using the predefined metrics, detailing its predictive capability on the development dataset. Aim 2: Model Validation and Optimization: To ensure the models are reliable and generalizable beyond the data on which they were trained, rigorous internal validation will be performed. This includes: 10-fold Cross-Validation: This technique will assess the model's stability by training and testing it on different subsets of the data, preventing overfitting. Bootstrap Resampling: This method will be used to generate multiple simulated datasets from the original sample to estimate the variability and stability of the model's performance metrics. The results from both development and validation phases will be systematically reported to ensure robustness. 3. Expected Outcomes and Deliverables The project is expected to produce several significant outcomes, which can be categorized into academic, practical, and policy-driven results: A Validated Prediction Model: The primary outcome will be a high-performing, internally validated machine learning model for predicting CVD risk in the Iranian population. The model will demonstrate its value by accurately identifying individuals at high risk, particularly those who might be missed by traditional models. Scientific and Academic Contributions: Research Protocol: A detailed protocol will be published and made publicly available (via OSF, protocols.io, etc.) to promote transparency and reproducibility. Peer-Reviewed Publications: The findings will be disseminated through publications in international scientific journals, detailing the model's development, validation, and potential clinical applications. Conference Presentations: The research will be presented at both national and international conferences, sharing methodological advancements and key findings with the scientific community. Technical Reports: Comprehensive technical documentation, including the Python and STATA code, will be generated for stakeholders and researchers in the cohort study network. Practical and Clinical Applications: The ultimate goal is to translate the research into a user-friendly digital tool. Development of a Digital Platform: The validated model will be implemented as a web-based calculator or a mobile application. This will provide healthcare professionals and even individuals with an accessible, real-time decision-support tool for personalized risk assessment. Policy and Public Health Impact: By providing a more accurate and locally relevant risk assessment tool, this research has the potential to: Enhance the accuracy and sustainability of cohort studies in Iran. Inform national health policies by demonstrating the critical role of dietary quality in CVD prevention. Improve the efficiency and effectiveness of precision public health strategies by allowing for more targeted prevention and management interventions.

// Source

View paper (DOI)Open access versionOpenAlexOpen Science FrameworkPublished 2026-08-03

Authors: Somaye Norouzi, Reza Chaman, Seyed Alireza Javadinia, Fereshteh Ghorat, Farnoosh Bakhshimoghadam