Achieving faithful explainability in feedforward neural networks through accurately computed feature attribution.

Journal: Neural networks : the official journal of the International Neural Network Society
Published Date:

Abstract

The rapid advancements in machine learning have led to the deployment of complex models in critical domains such as healthcare, finance, and autonomous systems. Despite their remarkable predictive performance, the opaque nature of these models presents significant challenges for interpretability, which is essential for trust, accountability, and regulatory compliance. Explainable Artificial Intelligence (XAI) has emerged as a crucial field addressing these challenges by making black-box models more transparent. In this paper, we propose a novel model-specific local post-hoc explanation method for feedforward neural networks (FNNs) built on a solid mathematical foundation. Our approach enables the exact computation of input feature attributions for individual predictions and achieves perfect fidelity in replicating model behavior. These two properties, combined with competitive computational efficiency, demonstrate the superior performance of the proposed method compared to state-of-the-art XAI techniques. We validate the method through extensive experiments, showing its versatility across diverse types of problems. This work enhances interpretability and trust in AI systems by providing a reliable explanation framework applicable across a wide range of scenarios modeled with FNNs.

Authors

Keywords

No keywords available for this article.