Improving breast cancer diagnosis accuracy using effective feature selection and machine learning
Abstract
Breast cancer (BC) is one of the leading causes of death among women worldwide. Early detection is crucial for saving lives and reducing mortality rates. However, the diagnostic process can be time-consuming due to the large number of features involved. Additionally, while various patient attributes are recorded in cancer datasets, not all of them are relevant for predicting the disease. To reduce costs, save time, and enhance the accuracy of BC diagnosis, this paper introduces a Feature selection (FS) algorithm based on the Correlation feature subset (CFS) method, combined with Logistic model tree (LMT). The main objectif is to identify the most relevant features for BC diagnosis. The proposed approach was evaluated using the Wisconsin diagnostic breast cancer (WDBC) dataset, available from the University of California, Irvine (UCI) repository. Without CFS based FS, using all 30 features, the highest accuracy achieved was 98.25% with LMT classifier. However, with the CFS based FS algorithm, the accuracy increased to 99.42%, while the number of features was reduced from 30 to 11. These findings demonstrate that the proposed CFS-based FS and LMT classifier improves diagnostic accuracy while selecting fewer, more effective features compared to the original dataset and these results outperform the existing methods. The results of this study could help develop a reliable clinical detection system, allowing experts to make more accurate and effective decisions in the future, making BC screening more affordable for patients by reducing the number of test parameters. Moreover, the proposed technology has the potential to be applied in detecting various types of cancers and other diseases.
// Source
Authors: Abderrahmane Ed-daoudy, Abdelouahed Ed-daoudy
Institutions: Chouaib Doukkali University, Abdelmalek Essaâdi University