New hybrid method enhances accuracy in ovarian cancer diagnosis
One of the primary obstacles in microarray-based cancer classification is the disproportionate number of features relative to patient samples, leading to model overfitting and reduced accuracy.
Ovarian cancer remains one of the most aggressive forms of cancer affecting women, often diagnosed at advanced stages due to a lack of early symptoms and effective screening methods. Early detection is crucial to improving survival rates, yet existing diagnostic tools often lack sensitivity and specificity. With advancements in data-driven healthcare, researchers have turned to computational approaches such as machine learning (ML) to enhance cancer identification. However, using ML on microarray data poses unique challenges due to its high dimensionality and class imbalance, where the number of gene expression features far exceeds the number of patient samples. Addressing these challenges requires specialized feature selection techniques and robust classification methods.
A recent study titled "Hybrid Methods to Identify Ovarian Cancer from Imbalanced High-Dimensional Microarray Data," conducted by Ni Kadek Emik Sapitri, Umu Sa'adah, and Nur Shofianah from Brawijaya University, Malang, Indonesia, explores novel hybrid approaches to ovarian cancer detection. Published in the IAES International Journal of Artificial Intelligence (IJ-AI), the research introduces two hybrid machine learning models that combine Infinite Feature Selection (IFS) with Classification and Regression Tree (CART) algorithms. These methods, termed SIFS-CART (Supervised Infinite Feature Selection with CART) and UIFS-CART (Unsupervised Infinite Feature Selection with CART), aim to improve classification accuracy while addressing high-dimensional and imbalanced data issues.
Overcoming data imbalance and high dimensionality
One of the primary obstacles in microarray-based cancer classification is the disproportionate number of features relative to patient samples, leading to model overfitting and reduced accuracy. The Infinite Feature Selection (IFS) algorithm, a graph-based feature selection technique, is designed to mitigate this problem by ranking features based on their importance. The study explores two variations: Supervised IFS (SIFS), which considers labeled data, and Unsupervised IFS (UIFS), which relies solely on feature interdependencies without labels. These methods aim to extract the most relevant genes associated with ovarian cancer, reducing computational complexity and enhancing predictive performance.
The dataset used in this study, OVA_ovary, consists of 10,937 gene expression features and 1,545 patient samples, making it highly imbalanced. To evaluate the effectiveness of their approach, the researchers tested three models:
- CART without feature selection (using all 10,935 features)
- SIFS-CART, selecting 1,000 most important features
- UIFS-CART, selecting 5,000 most important features
The study employed Balanced Accuracy (BA) as the primary evaluation metric to counteract biases introduced by imbalanced class distributions, ensuring that both cancerous and non-cancerous samples were weighted equally in performance assessments.
Evaluating model performance
The experimental results demonstrated that feature selection significantly enhanced model performance. SIFS-CART achieved the highest accuracy, outperforming both UIFS-CART and the CART baseline model. The key findings included:
- SIFS-CART achieved a balanced accuracy (BA) of 85.65% on training data and 83.23% on test data, using only 1,000 features.
- UIFS-CART, while an improvement over CART, performed less effectively, requiring 5,000 features to achieve a BA of 77.50% (training) and 75.74% (testing).
- The standard CART model, using all 10,935 features, achieved a lower BA of 84.44% (training) and 82.5% (testing), indicating that excessive features contributed to noise rather than meaningful classification improvements.
These results highlight the importance of feature selection in improving classification efficiency and accuracy. SIFS-CART not only outperformed the other methods in predictive accuracy but also demonstrated a more efficient use of features, reducing the computational burden while maintaining robust classification performance.
Implications and future directions
The study underscores the necessity of integrating feature selection techniques in microarray-based cancer classification to mitigate overfitting and improve interpretability. By reducing the feature set while retaining essential genetic markers, the SIFS-CART model offers a practical approach for ovarian cancer diagnosis. This method can be extended to other cancer types and high-dimensional biomedical datasets, opening new possibilities for AI-assisted medical diagnostics.
Future research could explore:
- Optimization of IFS parameters to further refine feature selection efficiency.
- Integration of deep learning models to enhance the identification of complex gene interactions.
- Application to real-world clinical datasets to validate the robustness of the approach in diverse patient populations.
In conclusion, the research by Sapitri et al. demonstrates that hybrid ML approaches combining Supervised Infinite Feature Selection (SIFS) and decision tree-based classification (CART) can significantly improve ovarian cancer detection. By leveraging data-driven insights, these methods pave the way for more accurate and scalable cancer diagnostic tools, ultimately contributing to better patient outcomes and early detection strategies.
- FIRST PUBLISHED IN:
- Devdiscourse
Google News