COUNTERFACTUAL EXPLANATIONS WITHOUT
OPENING THE BLACK BOX: AUTOMATED DECISIONS
AND THE GDPR
Sandra Wachter, Brent Mittelstadt, & Chris Russell
- Oxford Internet Institute, University of Oxford, 1 St. Giles, Oxford, OX1 3JS, UK and
The Alan Turing Institute, British Library, 96 Euston Road, London, NW1 2DB, UK. E-
mail: [email protected]. This work was supported by The Alan Turing
Institute under the EPSRC grant EP/N510129/1.
** Oxford Internet Institute, University of Oxford, 1 St. Giles, Oxford, OX1 3JS, UK,
The Alan Turing Institute, British Library, 96 Euston Road, London, NW1 2DB, UK,
Department of Science and Technology Studies, University College London, 22 Gordon
Square, London, WC1E 6BT, UK.
* The Alan Turing Institute, British Library, 96 Euston Road, London, NW1 2DB, UK,
Department of Electrical and Electronic Engineering, University of Surrey, Guildford,
GU2 7HX, UK.
2
COUNTERFACTUAL EXPLANATIONS
TABLE OF CONTENTS
I. Introduction ........................................................................................ 3
II. Counterfactuals ................................................................................. 5
A. Historic Context and The Problem of Knowledge ........................... 7
B. Explanations in A.I. and Machine Learning .................................. 10
C. Adversarial Perturbations and Counterfactual Explanations ....... 13
D. Causality and Fairness .................................................................. 15
III. Generating Counterfactuals ........................................................... 16
A. LSAT dataset .................................................................................. 18
B. Pima Diabetes Database ................................................................ 20
C. Causal Assumptions and Counterfactual Explanations ................ 21
IV. Advantages of Counterfactual Explanations .................................. 22
V. Counterfactual explanations and the GDPR ................................... 23
A. Explanations to understand decisions ........................................... 25
1. Broader possibilities with the right of access ............................ 32
2. Understanding through counterfactuals .................................... 34
B. Explanations to contest decisions .................................................. 35
1. Contesting through counterfactuals .......................................... 40
C. Explanations to alter future decisions ........................................... 42
Conclusion .......................................................................................... 43
Appendix 1: Simple Local Models as Explanations ............................. 48
Appendix 2: Example Transparency Infographic ................................. 51
3
COUNTERFACTUAL EXPLANATIONS
I. INTRODUCTION
There has been much discussion of the existence of a “right to
explanation” in the EU General Data Protection Regulation (“GDPR”),
and its merits and disadvantages.1 Attempts to implement a right to
explanation that opens the “black box” to provide insight into the internal
decision-making process of algorithms face four major legal and technical
barriers. First, a legally binding right to explanation does not exist in the
GDPR.2 Second, even if legally binding, the right would only apply in
limited cases (when a negative decision was solely automated and had
legal or other similar significant effects).3 Third, explaining the
functionality of complex algorithmic decision-making systems and their
rationale in specific cases is a technically challenging problem.4
Explanations may likewise offer little meaningful information to data
subjects, raising questions about their value.5 Finally, data controllers
have an interest in not sharing details of their algorithms to avoid
and
(June
Placing
SQUARESPACE
https://papers.ssrn.com/abstract=2964855
1 See, e.g., Sandra Wachter, Brent Mittelstadt & Luciano Floridi, Why a Right to
Explanation of Automated Decision-Making Does Not Exist in the General Data
Protection Regulation, 7 INT’L DATA PRIV. LAW 76, 79–90 (2017); Isak Mendoza & Lee
A. Bygrave, The Right Not to Be Subject to Automated Decisions Based on Profiling, in
EU INTERNET LAW: REGULATION AND ENFORCEMENT (Tatiani Synodinou et al. eds.,
[https://perma.cc/XV3T-G98W];
2017),
Lilian Edwards & Michael Veale, Slave to the Algorithm? Why a ‘Right to Explanation’
is Probably Not the Remedy You are Looking For, 16 DUKE L. TECH. REV. 18, 18–19
(2017); Tae Wan Kim & Bryan Routledge, Algorithmic Transparency, a Right to
2017),
Trust,
Explanation,
https://static1.squarespace.com/static/592ee286d482e908d35b8494/t/59552415579fb3
0c014cd06c/1498752022120/Algorithmic+transparency%2C+a+right+to+explanation+
and+trust+%28TWK%26BR%29.pdf [https://perma.cc/K53W-GVN2]; Gianclaudio
Malgieri & Giovanni Comandé, Why a Right to Legibility of Automated Decision-
Making Exists in the General Data Protection Regulation, 7 INT’L DATA PRIV. L. 243,
246–47 (2017); Bryce Goodman & Seth Flaxman, EU Regulations on Algorithmic
Decision-Making and a “Right to Explanation,” ARXIV:1606.08813, at 6–7 (2016),
http://arxiv.org/abs/1606.08813 [https://perma.cc/5ZTR-WG8R]; Andrew Selbst &
Julia Powles, Meaningful Information and the Right to Explanation, 7 INT’L DATA PRIV.
L. 233, 233–34 (2017).
2 Wachter, Mittelstadt & Floridi, supra note 1, at 79; Kim & Routledge, supra note 1, at
3.
3 Wachter, Mittelstadt & Floridi, supra note 1, at 78.
4 See, e.g., Wachter, Mittelstadt & Floridi, supra note 1, at 77; Edwards & Veale, supra
note 1, at 22; Joshua A. Kroll et al., Accountable Algorithms, 165 U. PA. L. REV. 633,
638 (2016); Tal Zarsky, Transparent Predictions, 2013 U. ILL. L. REV. 1503, 1519–20
(2013).
5 Jenna Burrell, How the Machine “Thinks:” Understanding Opacity in Machine
Learning Algorithms, BIG DATA & SOC., Jan.–June 2016, at 5; Kroll et al., supra note 4,
at 638.
4
COUNTERFACTUAL EXPLANATIONS
disclosing trade secrets, violating the rights and freedoms of others (e.g.
privacy), and allowing data subjects to game or manipulate the decision-
making system.6
Despite these difficulties, the social and ethical value (and perhaps
responsibility) of offering explanations to affected data subjects remains
unaffected. One significant point has been neglected in this discussion.
An explanation of automated decisions, both as envisioned by the GDPR
and in general, does not necessarily hinge on the general public
understanding of how algorithmic systems function. Even though such
interpretability is of great importance and should be pursued, explanations
can, in principle, be offered without opening the “black box.” Looking at
explanations as a means to help a data subject act rather than merely
understand, one could gauge the scope and content of explanations
according to the specific goal or action they are intended to support.
Explanations can serve many purposes. To investigate the
potential scope of explanations, it seems reasonable to start from the
perspective of the data subject, which is the natural person whose data is
being collected and evaluated. We propose three aims for explanations to
assist data subjects: (1) to inform and help the subject understand why a
particular decision was reached, (2) to provide grounds to contest adverse
decisions, and (3) to understand what could be changed to receive a
desired result in the future, based on the current decision-making model.
Publicité
As we show, the GDPR offers little support to achieve any of these aims.
However, none hinge on explaining the internal logic of automated
decision-making systems.
Building trust is essential to increase societal acceptance of
algorithmic decision-making. As a solution to close current gaps in
transparency and accountability that undermine trust between data
controllers and data subjects,7 we propose to move beyond the limitations
6 Burrell, supra note 5, at 3; Brenda Reddix-Smalls, Credit Scoring and Trade Secrecy:
An Algorithmic Quagmire or How the Lack of Transparency in Complex Financial
Models Scuttled the Finance Market, 12 U.C. DAVIS BUS. L.J. 87, 94 (2011); Mike
Ananny & Kate Crawford, Seeing without knowing: Limitations of the Transparency
Ideal and its Application to Algorithmic Accountability, NEW MEDIA & SOC., 2016, at 8,
http://journals.sagepub.com/doi/full/10.1177/1461444816676645
[https://perma.cc/3HF6-G9DS]; Roger A. Ford & W. Nicholson Price II, Privacy and
Accountability in Black-Box Medicine, 23 MICH. TELECOMM. TECH. REV. 1, 3 (2016);
Frank A. Pasquale, Restoring Transparency to Automated Authority, 9 J. TELECOMM.
HIGH TECH. L. 235, 237 (2011).
7 Wachter, Mittelstadt & Floridi, supra note 1, at 78; Mendoza & Bygrave, supra note
1, at 97.
5
COUNTERFACTUAL EXPLANATIONS
of the GDPR. We argue that counterfactuals should be used as a means to
provide explanations for individual decisions.
this paper, we present
Unconditional counterfactual explanations should be given for
positive and negative automated decisions, regardless of whether the
decisions are solely (as opposed to predominantly) automated or produce
legal or other significant effects. This approach provides data subjects
with meaningful explanations to understand a given decision, grounds to
contest it, and advice on how the data subject can change his or her
behaviour or situation to possibly receive a desired decision (e.g. loan
approval) in the future without facing the severely limited applicability
imposed by the GDPR’s definition of automated individual decision-
making.8
In
the concept of unconditional
counterfactual explanations as a novel type of explanation of automated
decisions that overcomes many challenges facing current work on
algorithmic interpretability and accountability. We situate counterfactuals
in the philosophical history of knowledge, as well as historical and
modern research on interpretability and fairness in machine learning.
Based on the potential advantages offered to data subjects by
counterfactual explanations, we then assess their alignment with the
GDPR’s numerous provisions concerning automated decision-making.
Specifically, we examine whether the GDPR offers support for
explanations that aim to help data subjects understand the scope of
automated decision-making as well as the rationale of specific decisions,
explanations to contest decisions, and explanations that offer guidance on
how data subjects can change their behaviour to receive a desired result.
We conclude that unconditional counterfactual explanations can bridge
the gap between the interests of data subjects and data controllers that
otherwise acts as a barrier to a legally binding right to explanation.
II. COUNTERFACTUALS
Counterfactual explanations take a similar form to the statement:
“You were denied a loan because your annual income was £30,000. If
your income had been £45,000, you would have been offered a loan.”
8 Wachter, Mittelstadt & Floridi, supra note 1, at 87–88; Mendoza & Bygrave, supra
note 1, at 83; Edwards & Veale, supra note 1, at 22.
6
COUNTERFACTUAL EXPLANATIONS
Here the statement of decision is followed by a counterfactual, or
statement of how the world would have to be different for a desirable
outcome to occur. Multiple counterfactuals are possible, as multiple
desirable outcomes can exist, and there may be several ways to achieve
any of these outcomes. The concept of the “closest possible world,” or the
smallest change to the world that can be made to obtain a desirable
outcome, is key throughout the discussion of counterfactuals. In many
situations, providing several explanations covering a range of diverse
counterfactuals corresponding to relevant or informative “close possible
worlds” rather than “the closest possible world” may be more helpful.
Knowing the smallest possible change to a variable or set of variables to
arrive at a different outcome may not always be the most helpful type of
counterfactual. Rather, relevance will depend also upon other case-
specific factors, such as the mutability of a variable or real world
probability of a change.9
In the existing literature, “explanation” typically refers to an
attempt to convey the internal state or logic of an algorithm that leads to
a decision.10 In contrast, counterfactuals describe a dependency on the
external facts that led to that decision. This is a crucial distinction. In
modern machine learning, the internal state of the algorithm can consist
of millions of variables intricately connected in a large web of dependent
behaviours.11 Conveying this state to a layperson in a way that allows
them to reason about the behaviour of an algorithm is extremely
challenging.12
The machine learning and legal communities have both taken
relatively restricted views on what passes for an explanation. The machine
learning community has been primarily concerned with debugging13 and
conveying approximations of algorithms that programmers or researchers
9 See infra, Section II.A.
10 See Burrell, supra note 5, at 1.
11 See, e.g., Kaiming He et al., Deep Residual Learning for Image Recognition, in
PROCEEDINGS OF THE IEEE CONFERENCE ON COMPUTER VISION AND PATTERN
RECOGNITION 770–78 (2016).
12 See Burrell, supra note 5, at 1; Zachary C. Lipton, The Mythos of Model
Interpretability, in 2016 WORKSHOP ON HUMAN INTERPRETABILITY IN MACHINE
96,
LEARNING
http://zacklipton.com/media/papers/mythos_model_interpretability_lipton2016.pdf
[https://perma.cc/4JVZ-7T6D].
13 Osbert Bastani, Carolyn Kim & Hamsa Bastani, Interpretability via Model Extraction,
AʀXɪᴠ:1706.09773,
https://arxiv.org/pdf/1611.07450.pdf
[https://perma.cc/8J3J-RE2T].
(2017),
at
1
7
COUNTERFACTUAL EXPLANATIONS
could use to understand which features are important14 while law and
ethics scholars have been more concerned with understanding the internal
logic of decisions as a means to assess their lawfulness (e.g. prevent
discriminatory outcomes), contest
increase accountability
generally, and clarify liability.15
them,
As such, the proposal made here for counterfactuals as
explanations lies outside of the taxonomies of explanations proposed
previously in machine learning, legal, and ethical literature. In contrast,
Publicité
as we discuss in the next section, analytic philosophy has taken a much
broader view of knowledge and how counterfactuals can be used as
justifications of beliefs.16
A. HISTORIC CONTEXT AND THE PROBLEM OF
KNOWLEDGE
Analytic Philosophy has a long history of analysing the necessary
conditions for propositional knowledge.17 Expressions of the type “S
1
at
and
at 1
DECISION-MAKING
Saliency Maps,
14 Marco Tulio Ribeiro, Sameer Singh & Carlos Guestrin, Why Should I Trust You?:
Explaining the Predictions of Any Classifier, in PROCEEDINGS OF THE 22ND ACM
SIGKDD INTERNATIONAL CONFERENCE ON KNOWLEDGE DISCOVERY AND DATA
MINING 1135 (2016); Ramprasaath R. Selvaraju et al., Grad-CAM: Why Did You Say
, https://arxiv.org/abs/1611.07450
(2016)
That?, ARXIV:1611.07450,
[https://perma.cc/AA8F-45XJ]; Karen Simonyan, Andrea Vedaldi & Andrew
Zisserman, Deep inside convolutional networks: Visualising Image Classification
Models
(2013),
ARXIV:1312.6034,
https://arxiv.org/abs/1312.6034 [https://perma.cc/Y85R-X9UE].
15 See, e.g., Finale Doshi-Velez et al., Accountability of AI Under the Law: The Role of
Explanation, ARXIV:1711.01134, at 1 (2017); Finale Doshi-Velez, Ryan Budish &
Mason Kortz, The Role of Explanation in Algorithmic Trust, TRUSTWORTHY
http://trustworthy-
ALGORITHMIC
algorithms.org/whitepapers/Finale%20Doshi-Velez.pdf [https://perma.cc/4L88-V58A];
Mireille Hildebrandt, The Dawn of a Critical Transparency Right for the Profiling Era,
in DIGITAL ENLIGHTENMENT YEARBOOK 2012 41 (Jacques Bus et al. eds., 2012); Tim
Miller, Explanation in Artificial Intelligence: Insights from the Social Sciences,
ARXIV:1706.07269, at 3 (2017); Pasquale, supra note 6, at 236; Danielle Keats Citron
& Frank A. Pasquale, The Scored Society: Due Process for Automated Predictions, 89
WASH. L. REV.
https://papers.ssrn.com/abstract=2376209
[https://perma.cc/9CXY-DBTN]; Tal Zarsky, Transparent Predictions, 2013 U. Iʟʟ. L.
(2013),
Rᴇᴠ.
https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2324240 [https://perma.cc/F8FC-
YDJG]; Tal Zarsky, The Trouble with Algorithmic Decisions: An Analytic Road Map to
Examine Efficiency and Fairness in Automated and Opaque Decision Making, 41 SCI.
TECH. HUM. VALUES 118, 118–132 (2016).
16 See, e.g., DAVID LEWIS, COUNTERFACTUALS 1–4, 84–91 (1973); David Lewis,
Counterfactuals and Comparative Possibility, 2 J. PHIL. LOGIC 418, 418–446 (1973);
Peter Lipton, Contrastive Explanation, 27 ROYAL INST. PHIL. SUPP. 247, 247 (1990);
ROBERT NOZICK, PHILOSOPHICAL EXPLANATIONS 172–74 (1981).
17 See generally ALFRED JULES AYER, THE PROBLEM OF KNOWLEDGE (1956).
1506–09
(2014),
1503,
6–7
2,
1,
8
COUNTERFACTUAL EXPLANATIONS
knows that p” constitute knowledge, where S refers to the knowing
subject, and p to the proposition that is known. Traditional approaches,
which conceive of knowledge as “justified true belief,” conceive of three
necessary conditions for knowledge: truth, belief, and justification.18
According to this tripartite approach, in order to know something, it is not
enough to simply believe that something is true: rather, you must also
have a good reason for believing it.19 The relevance of this approach
comes from the observation that this form of justification of beliefs can
serve as a type of explanation,20 as it is fundamentally a reason that a
belief is held and therefore serves as an answer to the question, “Why do
you believe X?” Understanding the different forms these justifications can
take opens the door to a broader class of explanations than previously
encountered in interpretability research.
Although influential, “justified true belief” has faced much
criticism21 and inspired substantial analysis of modifications to this
tripartite approach as well as proposals for additional necessary
conditions for a proposition to constitute knowledge.22 Modal conditions,
including safety23 and sensitivity,24 have been proposed as necessary
additions to the tripartite built on counterfactual relations.25
Sosa26 as well as Ichikawa and Steup27 define sensitivity as:
If p were false, S would not believe that p.
18 See generally Edmund L. Gettier, Is Justified True Belief Knowledge?, 23 ANALYSIS
121 (1963); Julien Dutant, The Legend of the Justified True Belief Analysis, 29 PHIL.
PERSP. 95 (2015).
19 See Gettier, supra note 18, at 121.
20 See NOZICK, supra note 16, at 174.
21 See, e.g., Dutant, supra note 18, at 95; Mark Kaplan, It’s Not What You Know that
Counts, 82 J. PHIL. 350, 350 (1985).
22 Jonathan Ichikawa & Matthias Steup, The Analysis of Knowledge, in STANFORD
ENCYCLOPEDIA
ed.),
PHILOSOPHY
https://plato.stanford.edu/archives/fall2017/entries/knowledge-analysis/
[https://perma.cc/6CXB-FTJV].
23 See Ernest Sosa, How to Defeat Opposition to Moore, 13 PHIL. PERSP. 141, 141–43
(1999).
24 See Jonathan Ichikawa, Quantifiers, Knowledge, and Counterfactuals, 82 PHIL.
PHENOMENOLOGICAL RES. 287, 287 (2011); see also NOZICK, supra note 16, at 172–74.
25 See Ichikawa & Steup, supra note 22, Section 5 (reviewing these concepts and their
criticisms).
26 Sosa, supra note 23, at 141.
27 Ichikawa & Steup, supra note 22, Section 5.1.
ARCHIVE
2017
(Fall
OF
9
COUNTERFACTUAL EXPLANATIONS
Here, the statement “If p were false” is a counterfactual defining a
“possible world” close to the world in which p is true.28 The sensitivity
condition suggests that “in the nearest possible worlds in which not-p, the
subject does not believe that p.”29 Our notion of counterfactual
explanations hinges upon the related concept:
If q were false, S would not believe p.
We claim that in this case, q serves as an explanation of S’s belief in p,
inasmuch as S only holds belief p while q is true, and that changing q
would also cause S’s belief to change. A key point is that such statements
only describe S’s beliefs, which need not reflect reality.30 As such, these
statements can be made without knowledge of any causal relationship
Publicité
between q and p.
We define Counterfactual Explanations as statements taking the
form:
Score p was returned because variables V had values (v1,
v2,...) associated with them. If V instead had values (v1',
v2',...), and all other variables had remained constant, score
p' would have been returned.
While many such explanations are possible, an ideal counterfactual
explanation would alter values as little as possible and represent a closest
world under which score p' is returned instead of p. The notion of a
“closest possible world” is thus implicit in our definition.
Our version of counterfactuals perhaps most resembles a
structural equations approach in execution by identifying alterations to
variables. This approach is more similar to Pearl’s “mini-surgeries”31 than
Lewis’ “miracles.”32 In any case, our approach does not rely on
knowledge of the causal structure of the world,33 or suggest which
context-dependent metric of distance between worlds is preferable to
establish causality.34 In many situations, it will be more informative to
provide a diverse set of counterfactual explanations, corresponding to
different choices of nearby possible worlds for which the counterfactual
28 See LEWIS, supra note 16, at 1–4.
29 Ichikawa & Steup, supra note 22, Section 5.1.
30 For example, S could believe that a person is inherently more trustworthy (p) because
they are a Capricorn (q).
31 See JUDEA PEARL, CAUSATION 223–24 (2000).
32 LEWIS, supra note 16, at 47–48.
33 See infra, Section II.D.
34 See Boris Kment, Counterfactuals and Explanation, 115 MIND 261, 261–309 (2006).
10
COUNTERFACTUAL EXPLANATIONS
holds or a preferred outcome is delivered, rather than a theoretically ideal
counterfactual describing the “closest possible world” according to a
preferred distance metric.35 Case-specific considerations will be relevant
to the choice of distance metric and a “sufficient” and “relevant” set of
counterfactual explanations. Such considerations may include the
capabilities of the individual concerned, sensitivity, mutability of the
variables involved in a decision, and ethical or legal requirements for
disclosure.36
Similarly, counterfactuals that describe changes to multiple
variables within the model can be provided. These would represent
possible futures brought about by changes
individual’s
circumstances. As an example, the impact of changes in income could be
calculated in combination with changes to career, thereby ensuring the
counterfactual represents a realistic possible world.
the
to
B. EXPLANATIONS IN A.I. AND MACHINE LEARNING
Much of the early work in A.I. on explaining the decisions made by expert
or rule-based systems focused on classes of explanation closely related to
counterfactuals. For example, Gregor and Benbasat37 offer the following
example of what they call a type 1 explanation:
Q: Why is a tax cut appropriate?
A: Because a tax cut’s preconditions are high inflation and
trade deficits, and current conditions include these factors.
35 The merits of different metrics of distance between possible worlds have long been
debated in philosophy without the emergence of consensus. Meaningfully addressing
this debate goes beyond the scope of this paper which proposes a method for
counterfactual explanations, but will be explored in future work. For further discussion
of distance metrics and counterfactuals, see LEWIS, supra note 16, at 8–15; Ernest W.
Adams, On the Rightness of Certain Counterfactuals, 74 PAC. PHIL. Q. 1, 1–8 (1993);
Kment, supra note 34, at 262.
36 A discussion of appropriate metrics for making these choices goes beyond the scope
of this paper, but will be addressed in future work. With that said, relevant philosophical
discussion can be found on determining relevance of possible causal or contrastive
explanations, counterfactuals, and distance metrics. See, e.g., Peter Lipton, Contrastive
Explanation, 27 ROYAL INST. PHIL. SUPP. 247, 254–65 (1990); Adams, supra note 35, at
1–8.
37 Shirley Gregor & Izak Benbasat, Explanations from Intelligent Systems: Theoretical
Foundations and Implications for Practice, 23 MIS Q. 497, 503 (1999).
11
COUNTERFACTUAL EXPLANATIONS
Buchanan and Shortliffe38 offer a similar example:
RULE009 IF:
1) The gram stain of the organism is gramneg, and
2) The morphology of the organism is coccus
THEN: There is strongly suggestive evidence (.8) that the
identity of the organism is Neisseria
As is typical in early A.I., questions we now recognise as hard such as
“How do we decide inflation is high?” or “Why are these the
preconditions of a tax cut?” are assumed to have been addressed by
humans, and are not discussed as part of the explanation.39 As such, the
explanations do not provide insight into what people in machine learning
think of as the internal logic of black box classifiers. In fact, the first
example can be rewritten as two diverse counterfactual statements:
“If inflation was lower, a tax cut would not be
recommended.”
“If there was no trade deficit, a tax cut would not be
recommended.”
While the second example is closely related to the counterfactual:40
“If the gram stain was negative or the morphology was
not coccus, the algorithm would not be confident that the
organism is Neisseria.”
important difference between
The most
these approaches and
counterfactuals is that counterfactuals continue functioning in an end-to-
end integrated approach. If the gram stain and morphology in the MYCIN
example were also determined by the algorithm, counterfactuals would
automatically return a close sample with a different classification, while
these early methods could not be applied to such involved scenarios.
38 BRUCE G. BUCHANAN & EDWARD D. SHORTLIFFE, RULE-BASED EXPERT SYSTEMS:
THE MYCIN EXPERIMENTS OF THE STANFORD HEURISTIC PROGRAMMING PROJECT 344
(1984).
39 See, e.g., Gregor & Benbasat, supra note 37, at 503.
40 However, they are not logically equivalent. The example from MYCIN differs in that
it is still possible that some samples that are either gram positive or have a different
morphology could still be classified as Neisseria.
12
COUNTERFACTUAL EXPLANATIONS
As focus has switched from A.I. and logic-based systems towards
machine learning tasks such as image recognition, the notion of an
explanation has come to refer to providing insight into the internal state
of an algorithm, or to human-understandable approximations of the
algorithm.41 As such, the most related machine learning work to these,
and to ours, is Martens and Provost.42 Uniquely among other works in
machine learning, it shares our interest in making interventions to alter
the outcome of classifier responses. However, the work is firmly linked
Publicité
to the problem of document classification, and the only interventions it
proposes involve the removal of words from documents to stop websites
from being classified as “adult.”43 The heuristic proposed cannot be easily
generalised to either continuous variables,44 or even the addition of words
to documents.
The majority of works in machine learning on explanations and
interpreting models concern themselves with generating simple models as
local approximations of decisions.45 Generally, the idea is to create a
simple human-understandable approximation of a decision-making
algorithm that accurately models the decision given the current inputs, but
may be arbitrarily bad for different inputs.46 However, there are numerous
difficulties with treating these approaches as explanations suitable for a
lay data subject.
In general, it is unclear if these models are interpretable by non-
experts. They make a three-way trade-off between the quality of the
approximation, the ease of understanding the function, and the size of the
domain for which the approximation is valid.47 As we show in Appendix
1, these local models can produce widely varying estimates of the
importance of variables even in simple scenarios such as the single
41 Ribeiro et al., supra note 14, at 1135–37.
42 David Martens & Foster Provost, Explaining Data-Driven Document Classifications,
38 MIS Q. 73, 73–74 (2013).
43 Id.
44 “Continuous variables” refers to variables whose assigned values are not restricted to
a small set of discrete values: such as present’ or not present’, but instead can take any
value in a given range. Measurements such as height, weight, or how bright a particular
pixel is in a photo, are often treated as continuous variables. See for example:
http://www.bbc.co.uk/schools/gcsebitesize/maths/statistics/samplinghirev1.shtml
45 Ribeiro et al., supra note 14, at 1135; Selvaraju et al., supra note 14, at 1–3; Simonyan
et al., supra note 14, at 1.
46 See Ribeiro et al., supra note 14, at 1143; Selvaraju et al., supra note 14, at 1–3;
Simonyan et al., supra note 14, at 1.
47 Bastani et al., supra note 13, at 1; Himabindu Lakkaraju et al., Interpretable &
Explorable Approximations of Black Box Models, AʀXɪᴠ:1707.01154, at 1 (2013),
https://arxiv.org/pdf/1707.01154.pdf [https://perma.cc/6JFE-N4YD].
13
COUNTERFACTUAL EXPLANATIONS
variable case, making it extremely difficult to reason about how a function
varies as the inputs change. Moreover, the utility of such approaches
outside of model debugging by expert programmers is unclear. Research
has yet to be conducted on how to convey the various limitations and
unreliabilities of these approaches to a lay audience in such a way that
they can make use of such explanations.
In contrast, counterfactual explanations are
intentionally
restricted. They are crafted in such a way as to provide a minimal amount
of information capable of altering a decision, and they do not require the
data subject to understand any of the internal logic of a model in order to
make use of it. The downside to this is that individual counterfactuals may
be overly restrictive. A single counterfactual may show how a decision is
based on certain data that is both correct and unable to be altered by the
data subject before future decisions, even if other data exist that could be
amended for a favourable outcome. This problem could be resolved by
offering multiple diverse counterfactual explanations to the data subject.
C. ADVERSARIAL PERTURBATIONS AND COUNTERFACTUAL
EXPLANATIONS
The techniques used to generate counterfactual explanations on
deep networks such as resnet48 are already widely studied in the machine
learning literature under the name of “Adversarial Perturbations.”49 In
these works, algorithms capable of computing counterfactuals are used to
confuse existing classifiers by generating a synthetic data point close to
an existing one such that the new synthetic data point is classified
differently than the original one.50
One strength of counterfactuals is that they can be efficiently and
effectively computed by applying standard techniques, even to cutting-
edge architectures. Some of the largest and deepest neural networks are
used in the field of computer vision, particularly in image labelling tasks
48 See He et al., supra note 11, at 770.
49 See Ian J. Goodfellow, Jonathon Shlens & Christian Szegedy, Explaining and
Harnessing Adversarial Examples, ARXIV:1412.6572
(2014),
https://arxiv.org/pdf/1412.6572.pdf
[https://perma.cc/64BR-WVE7]; Seyed-Mohsen
Moosavi-Dezfooli, Alhussein Fawzi & Pascal Frossard, Deepfool: A Simple and
Accurate Method to Fool Deep Neural Networks, in PROCEEDINGS OF THE IEEE
CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION 2574–82 (2016);
Christian Szegedy et al., Intriguing Properties of Neural Networks, ARXIV:1312.6199,
at 2 (2013) , https://arxiv.org/pdf/1312.6199.pdf [https://perma.cc/K37R-6NP2].
50 See Goodfellow, Shlens & Szegedy, supra note 49, at 1; Moosavi-Dezfooli, Fawzi &
Frossard, supra note 49, at 2574–82; Szegedy et al., supra note 49, at 2.
at
1
,
14
COUNTERFACTUAL EXPLANATIONS
such as ImageNet.51 These classifiers have been shown to be particularly
vulnerable to a type of attack referred to as “Adversarial Perturbation”
where small changes to a given image can result in the image being
assigned to an entirely different class. For example, DeepFool52 defines
an adverse perturbation of an image x, given a classifier, as the smallest
change to x such that the classification changes. Essentially, this is a
counterfactual by a different name. Finding a closest possible world to x
such that the classification changes is, under the right choice of distance
function, the same as finding the smallest change to x.
Importantly, none of
the standard works on Adversarial
Perturbations make use of appropriate distance functions, and the majority
of such approaches tend to favour making small changes to many
variables, instead of providing sparse human interpretable solutions that
modify only a few variables.53 Despite this, efficient computation of
counterfactuals and Adversarial Perturbations is made possible by virtue
of state-of-the-art algorithms being differentiable. Many optimisation
techniques proposed in the Adversarial Perturbation literature are directly
applicable to this problem, making counterfactual generation efficient.
One of the more challenging aspects of Adversarial Perturbations
is that these small perturbations of an image are barely human perceptible,
but result in drastically different classifier responses.54 Informally, this
appears to happen because the newly generated images do not lie in the
“space of real-images,” but slightly outside it.55 This phenomenon serves
as an important reminder that when computing counterfactuals by
searching for a close possible world, it is at least as important that the
solution found comes from a “possible world” as it is that it is close to the
starting example. Further research into how data from high-dimensional
and highly-structured spaces, such as natural images, can be characterised
is needed before counterfactuals can be reliably used as explanations in
these spaces.
51 See Jia Deng et al., Imagenet: A Large-Scale Hierarchical Image Database, in IEEE
CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION 24...