Multi-modal learning with incomplete data
Abstract
Multi-modal learning, in which diverse data types are integrated and analyzed together, has become a central area of research in artificial intelligence, driving major advances in a wide range of domains. However, in many practical situations, certain modalities or variables may be missing for part of the samples, leading to a limited performance or failure of conventional methods. This has given a rise to the field of multi-modal learning with incomplete data, an area that has grown rapidly due to its broad real-world applications. Despite this, the community still lacks standardized tools to effectively handle incomplete multi-modal data. To fill this gap, we developed iMML, a unified, user-friendly Python package with versatile methods designed for integrating, processing, and analyzing incomplete multi-modal data. Successful use cases in biomedicine, text analysis, and computer vision for diverse machine learning tasks show the potency of iMML for making the best use of modern datasets in complex real-world applications. The iMML package is available at https://github.com/ocbe-uio/imml with an extensive documentation at https://imml.readthedocs.io/. Incomplete multi-modal datasets pose a major challenge for real-world machine learning applications. Here, authors present iMML, a unified open-source Python package designed to analyze and integrate incomplete multi-modal datasets for diverse machine learning tasks.
// Source
Authors: Alberto López, John Zobolas, Tanguy Dumontier, Tero Aittokallio
Institutions: University of Helsinki, University of Oslo, Oslo University Hospital, Institute for Molecular Medicine Finland, Cancer Society of Finland