Publication details
- Journal: Machine Learning, vol. 115, Monday 20. July 2026
-
International Standard Numbers:
- Printed: 0885-6125
- Electronic: 1573-0565
- Links:
Abstract Today, the use of Artificial Intelligence (AI) is rapidly increasing in many areas of society. While model performance on various tasks continue to impress, it does so at the cost of increased model complexity, such that most state-of-the-art AI models are effectively black boxes. Where human-made decisions typically are accompanied by human-understandable explanations detailing the reasoning behind the decision, incorporating advanced AI as part of a decision-making process reduces the transparency of that process significantly. Yet, the ability to explain decisions is essential for there to be understanding and trust. As a response to this, Explainable Artificial Intelligence (XAI) has emerged as a field that aims to provide explanations of model behaviour. Methods categorised as post-hoc are designed to generate explanations for black box models after training, at no cost to model performance. In parallel with this, extensive work has been done in the field of causality to formalise the structure of human-understandable, causal explanations. This work presents a comprehensive literature review of the current state of the subfield of XAI that consist of causality-motivated post-hoc XAI methods. In order to clearly define causal XAI, a causal framework for categorising XAI is introduced, and three types of post-hoc XAI methods are identified: observational methods, internally causal methods and externally causal methods. Finally, externally causal XAI is argued a promising direction for reliable and understandable post-hoc XAI, with the ability to generate counterfactual explanations using a meaningful vocabulary, in line with the definition of counterfactual used in causal theory.