Explainable Pipelines for AI: Integrating Transparency into Data Engineering Workflows
DOI:
https://doi.org/10.70153/Keywords:
Explainable AI, Data Engineering, Transparency, Interpretability, Causal Inference, Data Lineage, Ethical AI, Feature Engineering, Data Provenance, Responsible AIAbstract
Artificial Intelligence (AI) systems are increasingly utilized in critical domains such as healthcare, finance, and governance, where transparency and accountability are essential. While explainable AI (XAI) research has primarily focused on model interpretability, the data engineering processes—including data ingestion, preprocessing, and feature engineering—remain largely opaque, posing challenges to trust, reproducibility, and ethical compliance. To bridge this gap, we propose an innovative Explainable Data Engineering (XDE) framework that integrates explainability throughout the entire data pipeline by leveraging techniques from explainable machine learning, causal inference, data provenance, and symbolic reasoning. We validate the framework using two real-world datasets: a breast cancer diagnosis dataset and a financial credit scoring dataset. In the healthcare setting, combining SHAP values with feature
lineage graphs enabled explanation of 98% of model decisions in terms of data transformations, while achieving a high classification accuracy of 93.5%, closely matching the traditional opaque pipeline. Medical experts rated the clarity of explanations highly, with an average score of 4.7 out of 5. For the financial dataset, the XDE pipeline successfully identified data drifts and anomalies overlooked by conventional methods, reducing false loan approvals by 12%. Narrative explanations facilitated compliance audits, enhancing stakeholder trust. Although the pipeline increased time-to-deployment by approximately 8%, it significantly reduced debugging time by 35%, improving maintainability. These results demonstrate that XDE effectively enhances transparency, auditability, and stakeholder confidence without sacrificing performance, offering a practical solution for responsible AI deployment through interpretable data pipelines.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.

