COUNTERFACTUAL EXPLANATIONS WITHOUT OPENING THE BLACK BOX: AUTOMATED DECISIONS AND THE GDPR

IEEE
Page 1 sur 52Lecteur de document UniversityLib

COUNTERFACTUAL EXPLANATIONS WITHOUT OPENING THE BLACK BOX: AUTOMATED DECISIONS AND THE GDPR

Artificial Intelligence, Machine Learning, Data Protection · notes

Browse all intelligence artificielle et données documents

COUNTERFACTUAL EXPLANATIONS WITHOUT

OPENING THE BLACK BOX: AUTOMATED DECISIONS

AND THE GDPR

Sandra Wachter, Brent Mittelstadt, & Chris Russell

  • Oxford Internet Institute, University of Oxford, 1 St. Giles, Oxford, OX1 3JS, UK and

The Alan Turing Institute, British Library, 96 Euston Road, London, NW1 2DB, UK. E-

mail: [email protected]. This work was supported by The Alan Turing

Institute under the EPSRC grant EP/N510129/1.

** Oxford Internet Institute, University of Oxford, 1 St. Giles, Oxford, OX1 3JS, UK,

The Alan Turing Institute, British Library, 96 Euston Road, London, NW1 2DB, UK,

Department of Science and Technology Studies, University College London, 22 Gordon

Square, London, WC1E 6BT, UK.

* The Alan Turing Institute, British Library, 96 Euston Road, London, NW1 2DB, UK,

Department of Electrical and Electronic Engineering, University of Surrey, Guildford,

GU2 7HX, UK.

2

COUNTERFACTUAL EXPLANATIONS

TABLE OF CONTENTS

I. Introduction ........................................................................................ 3

II. Counterfactuals ................................................................................. 5

A. Historic Context and The Problem of Knowledge ........................... 7

B. Explanations in A.I. and Machine Learning .................................. 10

C. Adversarial Perturbations and Counterfactual Explanations ....... 13

D. Causality and Fairness .................................................................. 15

III. Generating Counterfactuals ........................................................... 16

A. LSAT dataset .................................................................................. 18

B. Pima Diabetes Database ................................................................ 20

C. Causal Assumptions and Counterfactual Explanations ................ 21

IV. Advantages of Counterfactual Explanations .................................. 22

V. Counterfactual explanations and the GDPR ................................... 23

A. Explanations to understand decisions ........................................... 25

1. Broader possibilities with the right of access ............................ 32

2. Understanding through counterfactuals .................................... 34

B. Explanations to contest decisions .................................................. 35

1. Contesting through counterfactuals .......................................... 40

C. Explanations to alter future decisions ........................................... 42

Conclusion .......................................................................................... 43

Appendix 1: Simple Local Models as Explanations ............................. 48

Appendix 2: Example Transparency Infographic ................................. 51

3

COUNTERFACTUAL EXPLANATIONS

I. INTRODUCTION

There has been much discussion of the existence of a “right to

explanation” in the EU General Data Protection Regulation (“GDPR”),

and its merits and disadvantages.1 Attempts to implement a right to

explanation that opens the “black box” to provide insight into the internal

decision-making process of algorithms face four major legal and technical

barriers. First, a legally binding right to explanation does not exist in the

GDPR.2 Second, even if legally binding, the right would only apply in

limited cases (when a negative decision was solely automated and had

legal or other similar significant effects).3 Third, explaining the

functionality of complex algorithmic decision-making systems and their

rationale in specific cases is a technically challenging problem.4

Explanations may likewise offer little meaningful information to data

subjects, raising questions about their value.5 Finally, data controllers

have an interest in not sharing details of their algorithms to avoid

and

(June

Placing

SQUARESPACE

https://papers.ssrn.com/abstract=2964855

1 See, e.g., Sandra Wachter, Brent Mittelstadt & Luciano Floridi, Why a Right to

Explanation of Automated Decision-Making Does Not Exist in the General Data

Protection Regulation, 7 INT’L DATA PRIV. LAW 76, 79–90 (2017); Isak Mendoza & Lee

A. Bygrave, The Right Not to Be Subject to Automated Decisions Based on Profiling, in

EU INTERNET LAW: REGULATION AND ENFORCEMENT (Tatiani Synodinou et al. eds.,

[https://perma.cc/XV3T-G98W];

2017),

Lilian Edwards & Michael Veale, Slave to the Algorithm? Why a ‘Right to Explanation’

is Probably Not the Remedy You are Looking For, 16 DUKE L. TECH. REV. 18, 18–19

(2017); Tae Wan Kim & Bryan Routledge, Algorithmic Transparency, a Right to

2017),

Trust,

Explanation,

https://static1.squarespace.com/static/592ee286d482e908d35b8494/t/59552415579fb3

0c014cd06c/1498752022120/Algorithmic+transparency%2C+a+right+to+explanation+

and+trust+%28TWK%26BR%29.pdf [https://perma.cc/K53W-GVN2]; Gianclaudio

Malgieri & Giovanni Comandé, Why a Right to Legibility of Automated Decision-

Making Exists in the General Data Protection Regulation, 7 INT’L DATA PRIV. L. 243,

246–47 (2017); Bryce Goodman & Seth Flaxman, EU Regulations on Algorithmic

Decision-Making and a “Right to Explanation,” ARXIV:1606.08813, at 6–7 (2016),

http://arxiv.org/abs/1606.08813 [https://perma.cc/5ZTR-WG8R]; Andrew Selbst &

Julia Powles, Meaningful Information and the Right to Explanation, 7 INT’L DATA PRIV.

L. 233, 233–34 (2017).

2 Wachter, Mittelstadt & Floridi, supra note 1, at 79; Kim & Routledge, supra note 1, at

3.

3 Wachter, Mittelstadt & Floridi, supra note 1, at 78.

4 See, e.g., Wachter, Mittelstadt & Floridi, supra note 1, at 77; Edwards & Veale, supra

note 1, at 22; Joshua A. Kroll et al., Accountable Algorithms, 165 U. PA. L. REV. 633,

638 (2016); Tal Zarsky, Transparent Predictions, 2013 U. ILL. L. REV. 1503, 1519–20

(2013).

5 Jenna Burrell, How the Machine “Thinks:” Understanding Opacity in Machine

Learning Algorithms, BIG DATA & SOC., Jan.–June 2016, at 5; Kroll et al., supra note 4,

at 638.

4

COUNTERFACTUAL EXPLANATIONS

disclosing trade secrets, violating the rights and freedoms of others (e.g.

privacy), and allowing data subjects to game or manipulate the decision-

making system.6

Despite these difficulties, the social and ethical value (and perhaps

responsibility) of offering explanations to affected data subjects remains

unaffected. One significant point has been neglected in this discussion.

An explanation of automated decisions, both as envisioned by the GDPR

and in general, does not necessarily hinge on the general public

understanding of how algorithmic systems function. Even though such

interpretability is of great importance and should be pursued, explanations

can, in principle, be offered without opening the “black box.” Looking at

explanations as a means to help a data subject act rather than merely

understand, one could gauge the scope and content of explanations

according to the specific goal or action they are intended to support.

Explanations can serve many purposes. To investigate the

potential scope of explanations, it seems reasonable to start from the

perspective of the data subject, which is the natural person whose data is

being collected and evaluated. We propose three aims for explanations to

assist data subjects: (1) to inform and help the subject understand why a

particular decision was reached, (2) to provide grounds to contest adverse

decisions, and (3) to understand what could be changed to receive a

desired result in the future, based on the current decision-making model.

Advertisement

As we show, the GDPR offers little support to achieve any of these aims.

However, none hinge on explaining the internal logic of automated

decision-making systems.

Building trust is essential to increase societal acceptance of

algorithmic decision-making. As a solution to close current gaps in

transparency and accountability that undermine trust between data

controllers and data subjects,7 we propose to move beyond the limitations

6 Burrell, supra note 5, at 3; Brenda Reddix-Smalls, Credit Scoring and Trade Secrecy:

An Algorithmic Quagmire or How the Lack of Transparency in Complex Financial

Models Scuttled the Finance Market, 12 U.C. DAVIS BUS. L.J. 87, 94 (2011); Mike

Ananny & Kate Crawford, Seeing without knowing: Limitations of the Transparency

Ideal and its Application to Algorithmic Accountability, NEW MEDIA & SOC., 2016, at 8,

http://journals.sagepub.com/doi/full/10.1177/1461444816676645

[https://perma.cc/3HF6-G9DS]; Roger A. Ford & W. Nicholson Price II, Privacy and

Accountability in Black-Box Medicine, 23 MICH. TELECOMM. TECH. REV. 1, 3 (2016);

Frank A. Pasquale, Restoring Transparency to Automated Authority, 9 J. TELECOMM.

HIGH TECH. L. 235, 237 (2011).

7 Wachter, Mittelstadt & Floridi, supra note 1, at 78; Mendoza & Bygrave, supra note

1, at 97.

5

COUNTERFACTUAL EXPLANATIONS

of the GDPR. We argue that counterfactuals should be used as a means to

provide explanations for individual decisions.

this paper, we present

Unconditional counterfactual explanations should be given for

positive and negative automated decisions, regardless of whether the

decisions are solely (as opposed to predominantly) automated or produce

legal or other significant effects. This approach provides data subjects

with meaningful explanations to understand a given decision, grounds to

contest it, and advice on how the data subject can change his or her

behaviour or situation to possibly receive a desired decision (e.g. loan

approval) in the future without facing the severely limited applicability

imposed by the GDPR’s definition of automated individual decision-

making.8

In

the concept of unconditional

counterfactual explanations as a novel type of explanation of automated

decisions that overcomes many challenges facing current work on

algorithmic interpretability and accountability. We situate counterfactuals

in the philosophical history of knowledge, as well as historical and

modern research on interpretability and fairness in machine learning.

Based on the potential advantages offered to data subjects by

counterfactual explanations, we then assess their alignment with the

GDPR’s numerous provisions concerning automated decision-making.

Specifically, we examine whether the GDPR offers support for

explanations that aim to help data subjects understand the scope of

automated decision-making as well as the rationale of specific decisions,

explanations to contest decisions, and explanations that offer guidance on

how data subjects can change their behaviour to receive a desired result.

We conclude that unconditional counterfactual explanations can bridge

the gap between the interests of data subjects and data controllers that

otherwise acts as a barrier to a legally binding right to explanation.

II. COUNTERFACTUALS

Counterfactual explanations take a similar form to the statement:

“You were denied a loan because your annual income was £30,000. If

your income had been £45,000, you would have been offered a loan.”

8 Wachter, Mittelstadt & Floridi, supra note 1, at 87–88; Mendoza & Bygrave, supra

note 1, at 83; Edwards & Veale, supra note 1, at 22.

6

COUNTERFACTUAL EXPLANATIONS

Here the statement of decision is followed by a counterfactual, or

statement of how the world would have to be different for a desirable

outcome to occur. Multiple counterfactuals are possible, as multiple

desirable outcomes can exist, and there may be several ways to achieve

any of these outcomes. The concept of the “closest possible world,” or the

smallest change to the world that can be made to obtain a desirable

outcome, is key throughout the discussion of counterfactuals. In many

situations, providing several explanations covering a range of diverse

counterfactuals corresponding to relevant or informative “close possible

worlds” rather than “the closest possible world” may be more helpful.

Knowing the smallest possible change to a variable or set of variables to

arrive at a different outcome may not always be the most helpful type of

counterfactual. Rather, relevance will depend also upon other case-

specific factors, such as the mutability of a variable or real world

probability of a change.9

In the existing literature, “explanation” typically refers to an

attempt to convey the internal state or logic of an algorithm that leads to

a decision.10 In contrast, counterfactuals describe a dependency on the

external facts that led to that decision. This is a crucial distinction. In

modern machine learning, the internal state of the algorithm can consist

of millions of variables intricately connected in a large web of dependent

behaviours.11 Conveying this state to a layperson in a way that allows

them to reason about the behaviour of an algorithm is extremely

challenging.12

The machine learning and legal communities have both taken

relatively restricted views on what passes for an explanation. The machine

learning community has been primarily concerned with debugging13 and

conveying approximations of algorithms that programmers or researchers

9 See infra, Section II.A.

10 See Burrell, supra note 5, at 1.

11 See, e.g., Kaiming He et al., Deep Residual Learning for Image Recognition, in

PROCEEDINGS OF THE IEEE CONFERENCE ON COMPUTER VISION AND PATTERN

RECOGNITION 770–78 (2016).

12 See Burrell, supra note 5, at 1; Zachary C. Lipton, The Mythos of Model

Interpretability, in 2016 WORKSHOP ON HUMAN INTERPRETABILITY IN MACHINE

96,

LEARNING

http://zacklipton.com/media/papers/mythos_model_interpretability_lipton2016.pdf

[https://perma.cc/4JVZ-7T6D].

13 Osbert Bastani, Carolyn Kim & Hamsa Bastani, Interpretability via Model Extraction,

AʀXɪᴠ:1706.09773,

https://arxiv.org/pdf/1611.07450.pdf

[https://perma.cc/8J3J-RE2T].

(2017),

at

1

7

COUNTERFACTUAL EXPLANATIONS

could use to understand which features are important14 while law and

ethics scholars have been more concerned with understanding the internal

logic of decisions as a means to assess their lawfulness (e.g. prevent

discriminatory outcomes), contest

increase accountability

generally, and clarify liability.15

them,

As such, the proposal made here for counterfactuals as

explanations lies outside of the taxonomies of explanations proposed

previously in machine learning, legal, and ethical literature. In contrast,

Advertisement

as we discuss in the next section, analytic philosophy has taken a much

broader view of knowledge and how counterfactuals can be used as

justifications of beliefs.16

A. HISTORIC CONTEXT AND THE PROBLEM OF

KNOWLEDGE

Analytic Philosophy has a long history of analysing the necessary

conditions for propositional knowledge.17 Expressions of the type “S

1

at

and

at 1

DECISION-MAKING

Saliency Maps,

14 Marco Tulio Ribeiro, Sameer Singh & Carlos Guestrin, Why Should I Trust You?:

Explaining the Predictions of Any Classifier, in PROCEEDINGS OF THE 22ND ACM

SIGKDD INTERNATIONAL CONFERENCE ON KNOWLEDGE DISCOVERY AND DATA

MINING 1135 (2016); Ramprasaath R. Selvaraju et al., Grad-CAM: Why Did You Say

, https://arxiv.org/abs/1611.07450

(2016)

That?, ARXIV:1611.07450,

[https://perma.cc/AA8F-45XJ]; Karen Simonyan, Andrea Vedaldi & Andrew

Zisserman, Deep inside convolutional networks: Visualising Image Classification

Models

(2013),

ARXIV:1312.6034,

https://arxiv.org/abs/1312.6034 [https://perma.cc/Y85R-X9UE].

15 See, e.g., Finale Doshi-Velez et al., Accountability of AI Under the Law: The Role of

Explanation, ARXIV:1711.01134, at 1 (2017); Finale Doshi-Velez, Ryan Budish &

Mason Kortz, The Role of Explanation in Algorithmic Trust, TRUSTWORTHY

http://trustworthy-

ALGORITHMIC

algorithms.org/whitepapers/Finale%20Doshi-Velez.pdf [https://perma.cc/4L88-V58A];

Mireille Hildebrandt, The Dawn of a Critical Transparency Right for the Profiling Era,

in DIGITAL ENLIGHTENMENT YEARBOOK 2012 41 (Jacques Bus et al. eds., 2012); Tim

Miller, Explanation in Artificial Intelligence: Insights from the Social Sciences,

ARXIV:1706.07269, at 3 (2017); Pasquale, supra note 6, at 236; Danielle Keats Citron

& Frank A. Pasquale, The Scored Society: Due Process for Automated Predictions, 89

WASH. L. REV.

https://papers.ssrn.com/abstract=2376209

[https://perma.cc/9CXY-DBTN]; Tal Zarsky, Transparent Predictions, 2013 U. Iʟʟ. L.

(2013),

Rᴇᴠ.

https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2324240 [https://perma.cc/F8FC-

YDJG]; Tal Zarsky, The Trouble with Algorithmic Decisions: An Analytic Road Map to

Examine Efficiency and Fairness in Automated and Opaque Decision Making, 41 SCI.

TECH. HUM. VALUES 118, 118–132 (2016).

16 See, e.g., DAVID LEWIS, COUNTERFACTUALS 1–4, 84–91 (1973); David Lewis,

Counterfactuals and Comparative Possibility, 2 J. PHIL. LOGIC 418, 418–446 (1973);

Peter Lipton, Contrastive Explanation, 27 ROYAL INST. PHIL. SUPP. 247, 247 (1990);

ROBERT NOZICK, PHILOSOPHICAL EXPLANATIONS 172–74 (1981).

17 See generally ALFRED JULES AYER, THE PROBLEM OF KNOWLEDGE (1956).

1506–09

(2014),

1503,

6–7

2,

1,

8

COUNTERFACTUAL EXPLANATIONS

knows that p” constitute knowledge, where S refers to the knowing

subject, and p to the proposition that is known. Traditional approaches,

which conceive of knowledge as “justified true belief,” conceive of three

necessary conditions for knowledge: truth, belief, and justification.18

According to this tripartite approach, in order to know something, it is not

enough to simply believe that something is true: rather, you must also

have a good reason for believing it.19 The relevance of this approach

comes from the observation that this form of justification of beliefs can

serve as a type of explanation,20 as it is fundamentally a reason that a

belief is held and therefore serves as an answer to the question, “Why do

you believe X?” Understanding the different forms these justifications can

take opens the door to a broader class of explanations than previously

encountered in interpretability research.

Although influential, “justified true belief” has faced much

criticism21 and inspired substantial analysis of modifications to this

tripartite approach as well as proposals for additional necessary

conditions for a proposition to constitute knowledge.22 Modal conditions,

including safety23 and sensitivity,24 have been proposed as necessary

additions to the tripartite built on counterfactual relations.25

Sosa26 as well as Ichikawa and Steup27 define sensitivity as:

If p were false, S would not believe that p.

18 See generally Edmund L. Gettier, Is Justified True Belief Knowledge?, 23 ANALYSIS

121 (1963); Julien Dutant, The Legend of the Justified True Belief Analysis, 29 PHIL.

PERSP. 95 (2015).

19 See Gettier, supra note 18, at 121.

20 See NOZICK, supra note 16, at 174.

21 See, e.g., Dutant, supra note 18, at 95; Mark Kaplan, It’s Not What You Know that

Counts, 82 J. PHIL. 350, 350 (1985).

22 Jonathan Ichikawa & Matthias Steup, The Analysis of Knowledge, in STANFORD

ENCYCLOPEDIA

ed.),

PHILOSOPHY

https://plato.stanford.edu/archives/fall2017/entries/knowledge-analysis/

[https://perma.cc/6CXB-FTJV].

23 See Ernest Sosa, How to Defeat Opposition to Moore, 13 PHIL. PERSP. 141, 141–43

(1999).

24 See Jonathan Ichikawa, Quantifiers, Knowledge, and Counterfactuals, 82 PHIL.

PHENOMENOLOGICAL RES. 287, 287 (2011); see also NOZICK, supra note 16, at 172–74.

25 See Ichikawa & Steup, supra note 22, Section 5 (reviewing these concepts and their

criticisms).

26 Sosa, supra note 23, at 141.

27 Ichikawa & Steup, supra note 22, Section 5.1.

ARCHIVE

2017

(Fall

OF

9

COUNTERFACTUAL EXPLANATIONS

Here, the statement “If p were false” is a counterfactual defining a

“possible world” close to the world in which p is true.28 The sensitivity

condition suggests that “in the nearest possible worlds in which not-p, the

subject does not believe that p.”29 Our notion of counterfactual

explanations hinges upon the related concept:

If q were false, S would not believe p.

We claim that in this case, q serves as an explanation of S’s belief in p,

inasmuch as S only holds belief p while q is true, and that changing q

would also cause S’s belief to change. A key point is that such statements

only describe S’s beliefs, which need not reflect reality.30 As such, these

statements can be made without knowledge of any causal relationship

Advertisement

between q and p.

We define Counterfactual Explanations as statements taking the

form:

Score p was returned because variables V had values (v1,

v2,...) associated with them. If V instead had values (v1',

v2',...), and all other variables had remained constant, score

p' would have been returned.

While many such explanations are possible, an ideal counterfactual

explanation would alter values as little as possible and represent a closest

world under which score p' is returned instead of p. The notion of a

“closest possible world” is thus implicit in our definition.

Our version of counterfactuals perhaps most resembles a

structural equations approach in execution by identifying alterations to

variables. This approach is more similar to Pearl’s “mini-surgeries”31 than

Lewis’ “miracles.”32 In any case, our approach does not rely on

knowledge of the causal structure of the world,33 or suggest which

context-dependent metric of distance between worlds is preferable to

establish causality.34 In many situations, it will be more informative to

provide a diverse set of counterfactual explanations, corresponding to

different choices of nearby possible worlds for which the counterfactual

28 See LEWIS, supra note 16, at 1–4.

29 Ichikawa & Steup, supra note 22, Section 5.1.

30 For example, S could believe that a person is inherently more trustworthy (p) because

they are a Capricorn (q).

31 See JUDEA PEARL, CAUSATION 223–24 (2000).

32 LEWIS, supra note 16, at 47–48.

33 See infra, Section II.D.

34 See Boris Kment, Counterfactuals and Explanation, 115 MIND 261, 261–309 (2006).

10

COUNTERFACTUAL EXPLANATIONS

holds or a preferred outcome is delivered, rather than a theoretically ideal

counterfactual describing the “closest possible world” according to a

preferred distance metric.35 Case-specific considerations will be relevant

to the choice of distance metric and a “sufficient” and “relevant” set of

counterfactual explanations. Such considerations may include the

capabilities of the individual concerned, sensitivity, mutability of the

variables involved in a decision, and ethical or legal requirements for

disclosure.36

Similarly, counterfactuals that describe changes to multiple

variables within the model can be provided. These would represent

possible futures brought about by changes

individual’s

circumstances. As an example, the impact of changes in income could be

calculated in combination with changes to career, thereby ensuring the

counterfactual represents a realistic possible world.

the

to

B. EXPLANATIONS IN A.I. AND MACHINE LEARNING

Much of the early work in A.I. on explaining the decisions made by expert

or rule-based systems focused on classes of explanation closely related to

counterfactuals. For example, Gregor and Benbasat37 offer the following

example of what they call a type 1 explanation:

Q: Why is a tax cut appropriate?

A: Because a tax cut’s preconditions are high inflation and

trade deficits, and current conditions include these factors.

35 The merits of different metrics of distance between possible worlds have long been

debated in philosophy without the emergence of consensus. Meaningfully addressing

this debate goes beyond the scope of this paper which proposes a method for

counterfactual explanations, but will be explored in future work. For further discussion

of distance metrics and counterfactuals, see LEWIS, supra note 16, at 8–15; Ernest W.

Adams, On the Rightness of Certain Counterfactuals, 74 PAC. PHIL. Q. 1, 1–8 (1993);

Kment, supra note 34, at 262.

36 A discussion of appropriate metrics for making these choices goes beyond the scope

of this paper, but will be addressed in future work. With that said, relevant philosophical

discussion can be found on determining relevance of possible causal or contrastive

explanations, counterfactuals, and distance metrics. See, e.g., Peter Lipton, Contrastive

Explanation, 27 ROYAL INST. PHIL. SUPP. 247, 254–65 (1990); Adams, supra note 35, at

1–8.

37 Shirley Gregor & Izak Benbasat, Explanations from Intelligent Systems: Theoretical

Foundations and Implications for Practice, 23 MIS Q. 497, 503 (1999).

11

COUNTERFACTUAL EXPLANATIONS

Buchanan and Shortliffe38 offer a similar example:

RULE009 IF:

1) The gram stain of the organism is gramneg, and

2) The morphology of the organism is coccus

THEN: There is strongly suggestive evidence (.8) that the

identity of the organism is Neisseria

As is typical in early A.I., questions we now recognise as hard such as

“How do we decide inflation is high?” or “Why are these the

preconditions of a tax cut?” are assumed to have been addressed by

humans, and are not discussed as part of the explanation.39 As such, the

explanations do not provide insight into what people in machine learning

think of as the internal logic of black box classifiers. In fact, the first

example can be rewritten as two diverse counterfactual statements:

“If inflation was lower, a tax cut would not be

recommended.”

“If there was no trade deficit, a tax cut would not be

recommended.”

While the second example is closely related to the counterfactual:40

“If the gram stain was negative or the morphology was

not coccus, the algorithm would not be confident that the

organism is Neisseria.”

important difference between

The most

these approaches and

counterfactuals is that counterfactuals continue functioning in an end-to-

end integrated approach. If the gram stain and morphology in the MYCIN

example were also determined by the algorithm, counterfactuals would

automatically return a close sample with a different classification, while

these early methods could not be applied to such involved scenarios.

38 BRUCE G. BUCHANAN & EDWARD D. SHORTLIFFE, RULE-BASED EXPERT SYSTEMS:

THE MYCIN EXPERIMENTS OF THE STANFORD HEURISTIC PROGRAMMING PROJECT 344

(1984).

39 See, e.g., Gregor & Benbasat, supra note 37, at 503.

40 However, they are not logically equivalent. The example from MYCIN differs in that

it is still possible that some samples that are either gram positive or have a different

morphology could still be classified as Neisseria.

12

COUNTERFACTUAL EXPLANATIONS

As focus has switched from A.I. and logic-based systems towards

machine learning tasks such as image recognition, the notion of an

explanation has come to refer to providing insight into the internal state

of an algorithm, or to human-understandable approximations of the

algorithm.41 As such, the most related machine learning work to these,

and to ours, is Martens and Provost.42 Uniquely among other works in

machine learning, it shares our interest in making interventions to alter

the outcome of classifier responses. However, the work is firmly linked

Advertisement

to the problem of document classification, and the only interventions it

proposes involve the removal of words from documents to stop websites

from being classified as “adult.”43 The heuristic proposed cannot be easily

generalised to either continuous variables,44 or even the addition of words

to documents.

The majority of works in machine learning on explanations and

interpreting models concern themselves with generating simple models as

local approximations of decisions.45 Generally, the idea is to create a

simple human-understandable approximation of a decision-making

algorithm that accurately models the decision given the current inputs, but

may be arbitrarily bad for different inputs.46 However, there are numerous

difficulties with treating these approaches as explanations suitable for a

lay data subject.

In general, it is unclear if these models are interpretable by non-

experts. They make a three-way trade-off between the quality of the

approximation, the ease of understanding the function, and the size of the

domain for which the approximation is valid.47 As we show in Appendix

1, these local models can produce widely varying estimates of the

importance of variables even in simple scenarios such as the single

41 Ribeiro et al., supra note 14, at 1135–37.

42 David Martens & Foster Provost, Explaining Data-Driven Document Classifications,

38 MIS Q. 73, 73–74 (2013).

43 Id.

44 “Continuous variables” refers to variables whose assigned values are not restricted to

a small set of discrete values: such as present’ or not present’, but instead can take any

value in a given range. Measurements such as height, weight, or how bright a particular

pixel is in a photo, are often treated as continuous variables. See for example:

http://www.bbc.co.uk/schools/gcsebitesize/maths/statistics/samplinghirev1.shtml

45 Ribeiro et al., supra note 14, at 1135; Selvaraju et al., supra note 14, at 1–3; Simonyan

et al., supra note 14, at 1.

46 See Ribeiro et al., supra note 14, at 1143; Selvaraju et al., supra note 14, at 1–3;

Simonyan et al., supra note 14, at 1.

47 Bastani et al., supra note 13, at 1; Himabindu Lakkaraju et al., Interpretable &

Explorable Approximations of Black Box Models, AʀXɪᴠ:1707.01154, at 1 (2013),

https://arxiv.org/pdf/1707.01154.pdf [https://perma.cc/6JFE-N4YD].

13

COUNTERFACTUAL EXPLANATIONS

variable case, making it extremely difficult to reason about how a function

varies as the inputs change. Moreover, the utility of such approaches

outside of model debugging by expert programmers is unclear. Research

has yet to be conducted on how to convey the various limitations and

unreliabilities of these approaches to a lay audience in such a way that

they can make use of such explanations.

In contrast, counterfactual explanations are

intentionally

restricted. They are crafted in such a way as to provide a minimal amount

of information capable of altering a decision, and they do not require the

data subject to understand any of the internal logic of a model in order to

make use of it. The downside to this is that individual counterfactuals may

be overly restrictive. A single counterfactual may show how a decision is

based on certain data that is both correct and unable to be altered by the

data subject before future decisions, even if other data exist that could be

amended for a favourable outcome. This problem could be resolved by

offering multiple diverse counterfactual explanations to the data subject.

C. ADVERSARIAL PERTURBATIONS AND COUNTERFACTUAL

EXPLANATIONS

The techniques used to generate counterfactual explanations on

deep networks such as resnet48 are already widely studied in the machine

learning literature under the name of “Adversarial Perturbations.”49 In

these works, algorithms capable of computing counterfactuals are used to

confuse existing classifiers by generating a synthetic data point close to

an existing one such that the new synthetic data point is classified

differently than the original one.50

One strength of counterfactuals is that they can be efficiently and

effectively computed by applying standard techniques, even to cutting-

edge architectures. Some of the largest and deepest neural networks are

used in the field of computer vision, particularly in image labelling tasks

48 See He et al., supra note 11, at 770.

49 See Ian J. Goodfellow, Jonathon Shlens & Christian Szegedy, Explaining and

Harnessing Adversarial Examples, ARXIV:1412.6572

(2014),

https://arxiv.org/pdf/1412.6572.pdf

[https://perma.cc/64BR-WVE7]; Seyed-Mohsen

Moosavi-Dezfooli, Alhussein Fawzi & Pascal Frossard, Deepfool: A Simple and

Accurate Method to Fool Deep Neural Networks, in PROCEEDINGS OF THE IEEE

CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION 2574–82 (2016);

Christian Szegedy et al., Intriguing Properties of Neural Networks, ARXIV:1312.6199,

at 2 (2013) , https://arxiv.org/pdf/1312.6199.pdf [https://perma.cc/K37R-6NP2].

50 See Goodfellow, Shlens & Szegedy, supra note 49, at 1; Moosavi-Dezfooli, Fawzi &

Frossard, supra note 49, at 2574–82; Szegedy et al., supra note 49, at 2.

at

1

,

14

COUNTERFACTUAL EXPLANATIONS

such as ImageNet.51 These classifiers have been shown to be particularly

vulnerable to a type of attack referred to as “Adversarial Perturbation”

where small changes to a given image can result in the image being

assigned to an entirely different class. For example, DeepFool52 defines

an adverse perturbation of an image x, given a classifier, as the smallest

change to x such that the classification changes. Essentially, this is a

counterfactual by a different name. Finding a closest possible world to x

such that the classification changes is, under the right choice of distance

function, the same as finding the smallest change to x.

Importantly, none of

the standard works on Adversarial

Perturbations make use of appropriate distance functions, and the majority

of such approaches tend to favour making small changes to many

variables, instead of providing sparse human interpretable solutions that

modify only a few variables.53 Despite this, efficient computation of

counterfactuals and Adversarial Perturbations is made possible by virtue

of state-of-the-art algorithms being differentiable. Many optimisation

techniques proposed in the Adversarial Perturbation literature are directly

applicable to this problem, making counterfactual generation efficient.

One of the more challenging aspects of Adversarial Perturbations

is that these small perturbations of an image are barely human perceptible,

but result in drastically different classifier responses.54 Informally, this

appears to happen because the newly generated images do not lie in the

“space of real-images,” but slightly outside it.55 This phenomenon serves

as an important reminder that when computing counterfactuals by

searching for a close possible world, it is at least as important that the

solution found comes from a “possible world” as it is that it is close to the

starting example. Further research into how data from high-dimensional

and highly-structured spaces, such as natural images, can be characterised

is needed before counterfactuals can be reliably used as explanations in

these spaces.

51 See Jia Deng et al., Imagenet: A Large-Scale Hierarchical Image Database, in IEEE

CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION 24...