A survey of inevitability results in AI instability.
Journal:
Philosophical transactions. Series A, Mathematical, physical, and engineering sciences
Published Date:
Jul 16, 2026
Abstract
Over the past decade, there has been an explosion of activity in the design of algorithms for adversarially attacking artificial intelligence (AI) systems; especially in the context of image classification. For example, a carefully crafted perturbation to an image that is imperceptible to the human eye may cause a sophisticated convolutional neural network to change its classification. To deal with this vulnerability, algorithms that detect or defend against such attacks have also been proposed. It has been observed empirically that attackers have the upper hand. To explain these observations, various theoretical results have subsequently emerged. This article will review some recent rigorous results, emphasizing what assumptions they use, in terms of (i) the nature of the training data, (ii) the network architectures, and (iii) the information available to the attacker. We also discuss how the results may give guidelines for building more secure systems, and how this research area can inform the design of AI regulations. This article is part of the theme issue 'Safe, secure and robust AI for safety-critical systems'.
Authors
Keywords
No keywords available for this article.