AI & Computingpreprint2026-08-18

Agentic AI for Data Engineering: A Systematic Review of Autonomous Pipeline Construction, Optimization, and Governance

Open access0 citations

Abstract

Data engineering is increasingly cited as a proving ground for agentic AI, with surveys reporting daily AI tooluse among 82% of data professionals — yet most of that use remains tactical rather than autonomous, and only a small fraction of teams describe their data modeling practice as mature. This survey asks why, mapping current agentic AI research and tooling onto five data engineering pipeline stages — ingestion and schema mapping, transformation and code generation, data quality and anomaly detection, orchestration and self-healing, and governance and lineage — and assessing the maturity, evaluation rigor, and structural risk of each. We find capability is real but uneven: transformation and code generation is the most mature stage, with systems reporting end-to-end success rates above 90%, while governance and lineage is the least mature, with the strongest available benchmark showing that even a purpose-built architecture solves barely half of complex governance tasks correctly. Reviewing how these systems are evaluated, we identify three compounding gaps — single-run success bias, a documented drop in reliability from 60% to 25% once agents are tested across repeated runs rather than once, and the near-absence of cost, latency, and compliance as evaluation dimensions alongside correctness. We then synthesize six open structural challenges — context and semantic grounding, trust and hallucination propagation, ownership and accountability, cost and compute tradeoffs, legacy system integration, and security of autonomous data access — arguing these are interdependent rather than separable, with weak semantic grounding as the root cause underlying several of the others. We conclude that agentic AI is being adopted in data engineering faster than the semantic and governance infrastructure needed to make that adoption reliable is beingbuilt, and outline concrete directions — production-realistic multi-dimensional benchmarks, context engineering as first-class infrastructure, and independent replication of the field's most-cited results — for closing that gap

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-18

Authors: Shakeel Daroga