AI & Computingarticle2026-09-09

A Feature Model-Based Reference Architecture for Data Lake Ingestion: A Variability Management Approach

Open access0 citations

Abstract

The growth of big data ecosystems has shifted the classical paradigm of selective storage toward an approach that preserves large volumes of heterogeneous data for subsequent exploitation, thereby strengthening the adoption of repositories such as data warehouses and, particularly, data lakes. In this context, data ingestion from multiple sources, formats, and structures is a critical activity in implementing and operating these environments. However, it is often carried out in a highly ad hoc manner, with low levels of standardization and with variability managed informally. Beyond the operational complexity of ingestion itself, the variability in features such as source types, ingestion frequencies, transformation needs, and loading strategies constitutes an additional engineering problem that must be addressed systematically. This work tackles both issues in the context of a consulting firm involved in data migration projects to data lakes under governance constraints defined by clients in the BFSI sector. The goal is to formalize the data ingestion process through an architecture that provides technical, documentation, and training support for engineering teams, while also incorporating a variability management tool to model, analyze, and guide the configuration of ingestion solutions according to project-specific needs. In this way, the proposal seeks to reduce uncertainty, improve development quality, optimize resource utilization, and provide a more systematic treatment of variability in data ingestion projects. The proposal was evaluated through a structured survey answered by two cohorts totaling 29 respondents: an enterprise cohort of 15 practitioners (60% of the firm’s staff) and a prospective cohort of 14 engineering interns. For the enterprise cohort, the survey obtained average scores of 76 (individual) and 83.1 (role-averaged) for usability and 85.71 (individual) and 91.56 (role-averaged) for perceived quality; for the intern cohort, the corresponding individual averages were 70.18 for usability and 74.74 for perceived quality. In addition, a before/after comparison against the previous ad hoc workflow, covering objective engineering indicators, was conducted with both cohorts, providing task-based evidence that complements the perception-based results.

// Source

View paper (DOI)Open access versionOpenAlexApplied SciencesPublished 2026-09-09

Authors: Juan Lagos-Obando, Oscar Aguayo, Raúl Mazo

Institutions: Université de Bretagne Occidentale, Universidad de La Frontera, Laboratoire des Sciences et Techniques de l’Information de la Communication et de la Connaissance