Adversarial Robustness: Theory and Practice

Page 1 sur 49Lecteur de document UniversityLib

Adversarial Robustness: Theory and Practice

Machine Learning, Adversarial Examples, Security · notes

Browse all intelligence artificielle et données documents

Adversarial Robustness:

Theory and Practice

Zico Kolter

Aleksander Mądry

madry-lab.ml

Tutorial website:

adversarial-ml-tutorial.org

@zicokolter

@aleks_madry

Machine Learning: The Success Story

Image classification

Machine translation

Reinforcement Learning

Machine Learning: The Success Story

Is ML truly ready for

real-world deployment?

Can We Truly Rely on ML?

ImageNet: An ML Home Run

ILSVRC top-5 Error on ImageNet

AlexNet

30

25

20

15

10

5

0

2010

2011

2012

2013

2014

Human

2015

2016

2017

But what do these results really mean?

A Limitation of the (Supervised) ML Framework

Measure of performance:

Fraction of mistakes during testing

But: In reality, the distributions

we use ML on are NOT the ones

we train it on

Training

Inference

A Limitation of the (Supervised) ML Framework

=

Measure of performance:

Fraction of mistakes during testing

But: In reality, the distributions

we use ML on are NOT the ones

we train it on

Training

Inference

What can go wrong?

ML Predictions Are (Mostly) Accurate but Brittle

“pig” (91%)

noise (NOT random)

“airliner” (99%)

+ 0.005 x

=

[Szegedy Zaremba Sutskever Bruna Erhan Goodfellow Fergus 2013]

[Biggio Corona Maiorca Nelson Srndic Laskov Giacinto Roli 2013]

But also: [Dalvi Domingos Mausam Sanghai Verma 2004][Lowd Meek 2005]

[Globerson Roweis 2006][Kolcz Teo 2009][Barreno Nelson Rubinstein Joseph Tygar 2010]

[Biggio Fumera Roli 2010][Biggio Fumera Roli 2014][Srndic Laskov 2013]

ML Predictions Are (Mostly) Accurate but Brittle

[Kurakin Goodfellow Bengio 2017]

[Sharif Bhagavatula Bauer Reiter 2016]

[Athalye Engstrom Ilyas Kwok 2017]

[Eykholt Evtimov Fernandes Li Rahmati Xiao Prakash Kohno Song 2017]

ML Predictions Are (Mostly) Accurate but Brittle

[Fawzi Frossard 2015]

[Engstrom Tran Tsipras Schmidt M 2018]:

Rotation + Translation suffices to fool

state-of-the-art vision models

→ Data augmentation does not

seem to help here either

So: Brittleness of ML is a thing

Should we be worried?

Why Is This Brittleness of ML a Problem?

→ Security

[Carlini Wagner 2018]:

Voice commands that are

unintelligible to humans

[Sharif Bhagavatula Bauer Reiter 2016]:

Glasses that fool face recognition

Why Is This Brittleness of ML a Problem?

→ Security

→ Safety

https://www.youtube.com/watch?v=TIUU1xNqI8w

https://www.youtube.com/watch?v=_1MHGUC_BzQ

Why Is This Brittleness of ML a Problem?

→ Security

→ Safety

→ ML Alignment

Need to understand the

“failure modes” of ML

Is That It?

Training

Inference

Data poisoning

Adversarial Examples

(Deep) ML is ”data hungry”

→ Can’t afford to be too picky about

where we get the training data from

What can go wrong?

Data Poisoning

Goal: Maintain training accuracy but hamper generalization

Advertisement

Data Poisoning

Goal: Maintain training accuracy but hamper generalization

→ Fundamental problem

in “classic” ML (robust statistics)

→ But: seems less so in deep learning

→ Reason: Memorization?

Data Poisoning

classification of specific inputs

Goal: Maintain training accuracy but hamper generalization

→ Fundamental problem

in “classic” ML (robust statistics)

→ But: seems less so in deep learning

→ Reason: Memorization?

Is that it?

Data Poisoning

classification of specific inputs

Goal: Maintain training accuracy but hamper generalization

[Koh Liang 2017]: Can manipulate many

predictions with a single “poisoned” input

But: This gets (much) worse

“van”

“dog”

[Gu Dolan-Gavitt Garg 2017][Turner Tsipras M 2018]:

Can plant an undetectable backdoor that

gives an almost total control over the model

(To learn more about backdoor attacks:

See poster #148 on Wed [Tran Li M 2018])

Is That It?

Microsoft Azure (Language Services)

Google Cloud Vision API

!

t

u

p

n

I

O

u

t

p

u

t

Parameters "

Training

Inference

Deployment

Is That It?

Does limited access

give security?

In short: No

!

t

u

p

n

I

O

u

t

p

u

t

Parameters "

Data

Predictions

Training

Inference

Deployment

Black box attacks

Is That It?

Does limited access

give security?

Model stealing: “Reverse

engineer“ the model

[Tramer Zhang Juels Reiter Ristenpart 2016]

Black box attacks: Construct

adv. examples from queries

[Chen Zhang Sharma Yi Hsieh 2017][Bhagoji He Li

Song 2017][Ilyas Engstrom Athalye Lin 2017]

[Brendel Rauber Bethge 2017][Cheng Le Chen Yi

Zhang Hsieh 2018][Ilyas Engstrom M 2018]

Data

!

t

u

p

n

I

O

u

t

p

u

t

Parameters "

Predictions

For more: See my talk on Friday

Training

Inference

Deployment

Black box attacks

Three commandments of Secure/Safe ML

I. Thou shall not train on data you don’t fully trust

(because of data poisoning)

II. Thou shall not let anyone use your model (or observe its

outputs) unless you completely trust them

(because of model stealing and black box attacks)

III. Thou shall not fully trust the predictions of your model

(because of adversarial examples)

Advertisement

Are we doomed?

(Is ML inherently not reliable?)

No: But we need to re-think how we do ML

(Think: adversarial aspects = stress-testing our solutions)

Towards Adversarially Robust Models

“pig” (91%)

“pig”

“airliner” (99%)

+ 0.005 x

=

Where Do Adversarial Examples Come From?

Differentiable

To get an adv. example

Goal of training:

Model Parameters

Input Correct Label

!"#$ %&'' $, ) , *

+

t

u

p

n

I

O

u

t

p

u

t

Parameters ,

Can use gradient descent

method to find good $

Where Do Adversarial Examples Come From?

Differentiable

To get an adv. example

Goal of training:

!"#$ %&'' (, # + $, +

,

t

u

p

n

I

O

u

t

p

u

t

Parameters -

Can use gradient descent

method to find good (

Where Do Adversarial Examples Come From?

Differentiable

To get an adv. example

Goal of training:

!"#$ %&'' (, # + $, +

,

t

u

p

n

I

O

u

t

p

u

t

Parameters -

Which $ are allowed?

Examples: $ that is small wrt

• ℓ/-norm

• Rotation and/or translation

• VGG feature perturbation

(add the perturbation you need here)

Can use gradient descent

This is an important question

method to find bad $

(that we put aside)

Still: We have to confront

(small) ℓ/-norm perturbations

Towards ML Models that Are Adv. Robust

[M Makelov Schmidt Tsipras Vladu 2018]

Key observation: Lack of adv. robustness is NOT at odds with

what we currently want our ML models to achieve

Standard generalization:

!(#,%)~( [*+,, -, ., / ]

Adversarially robust

But: Adversarial noise is a “needle in a haystack”

Towards ML Models that Are Adv. Robust

[M Makelov Schmidt Tsipras Vladu 2018]

Key observation: Lack of adv. robustness is NOT at odds with

what we currently want our ML models to achieve

Standard generalization:

Adversarially robust

!(#,%)~( [*+,

-∈/

0122 3, , + -, 5 ]

But: Adversarial noise is a “needle in a haystack”

Next: A deeper dive into the topic

→ Adversarial examples and verification (Zico)

→ Training adversarially robust models (Zico)

→ Adversarial robustness beyond security (Aleksander)

Adversarial Robustness Beyond Security

ML via Adversarial Robustness Lens

Overarching question:

How does adv. robust ML differ from “standard” ML?

Advertisement

!(#,%)~( [*+,, -, ., / ]

vs

!(#,%)~( [12.

3∈5

*+,, -, . + 3, / ]

(This goes beyond deep learning)

Do Robust Deep Networks Overfit?

Accuracy

100%

80%

60%

40%

20%

0%

0

10000

20000

30000

40000

50000

60000

70000

80000

Std Training

Do Robust Deep Networks Overfit?

Accuracy

100%

80%

60%

40%

20%

0%

0

10000

20000

30000

40000

50000

60000

70000

80000

Std Training

Std Evaluation

(small)

generalization gap

Do Robust Deep Networks Overfit?

Accuracy

100%

80%

60%

40%

20%

0%

0

10000

20000

30000

40000

50000

60000

70000

80000

Adv Trainining

Do Robust Deep Networks Overfit?

Accuracy

100%

80%

60%

40%

20%

0%

(large)

generalization gap

Regularization does not

seem to help either

0

10000

20000

30000

40000

50000

60000

70000

80000

Adv Evaluation

Adv Trainining

What’s going on?

Adv. Robust Generalization Needs More Data

Theorem [Schmidt Santurkar Tsipras Talwar M 2018]:

Sample complexity of adv. robust generalization can be

significantly larger than that of “standard” generalization

Specifically: There exists a d-dimensional distribution D s.t.:

$∗

→ A single sample is enough to get an accurate

classifier (P[correct] > 0.99)

→ But: Need ! " samples for better-than-chance

robust classifier

(More details: See spotlight + poster #31 on Tue)

−$∗

+$

−$

Does Being Robust Help “Standard” Generalization?

Data augmentation: An effective technique

to improve “standard” generalization

Adversarial training

=

An “ultimate” version of data augmentation?

(since we train on the ”most confusing” version of the training set)

Does adversarial training always improve

Advertisement

“standard” generalization?

Does Being Robust Help “Standard” Generalization?

Accuracy

100%

80%

60%

40%

20%

0%

0

10000

20000

30000

40000

50000

60000

70000

80000

Std Evaluation of Std Training

Does Being Robust Help “Standard” Generalization?

Accuracy

100%

80%

60%

40%

20%

0%

“standard”

performance gap

Where is this

(consistent) gap

coming from?

0

10000

20000

30000

40000

50000

60000

70000

80000

Std Eval of Adv. Training

Std Evaluation of Std Training

Does Being Robust Help “Standard” Generalization?

Theorem [Tsipras Santurkar Engstrom Turner M 2018]:

No “free lunch”: can exist a trade-off between accuracy and robustness

Basic intuition:

→ In standard training, all correlation is good correlation

→ If we want robustness, must avoid weakly correlated features

Strong (but not perfect)

correlation

aggregates to a very accurate (but non-robust!) “meta-feature”

Standard training: use all of

features, maximize accuracy

Weak correlation

Adversarial training: use only single robust

feature (at the expense of accuracy)

Adversarial Robustness is Not Free

→ Optimization during training more difficult

and models need to be larger

→ More training data might be required

[Schmidt Santurkar Tsipras Talwar M 2018]

+"

−"

→ Might need to lose on “standard” measures of performance

[Tsipras Santurkar Engstrom Turner M 2018] (Also see: [Bubeck Price Razenshteyn 2018])

But There Are (Unexpected?) Benefits Too

[Tsipras Santurkar Engstrom Turner M 2018]

Models become more semantically meaningful

Input

Gradient of

standard model

Gradient of

adv. robust model

But There Are (Unexpected?) Benefits Too

[Tsipras Santurkar Engstrom Turner M 2018]

Models become more semantically meaningful

“Bird”

“Primate”

“Bird”

“Primate”

Standard model

Adv. robust model

[Brock Donahue Simonyan 2018]

+ [Isola 2018]

Robust models → (restricted) GAN-like embeddings?

Conclusions

Towards (Adversarially) Robust ML

→ Algorithms: Faster robust training + verification [Xiao Tjeng Shafiullah M 2018],

smaller models, new architectures?

→ Theory: (Better) adv. robust generalization bounds,

new regularization techniques

→ Data: New datasets and more comprehensive set of perturbations

Major need: Embracing more of a worst-case mindset

→ Adaptive evaluation methodology + scaling up verification

(robust-ml.org)

More Broadly

Next frontier:

Building ML one can truly rely on

→ Will lead to ML that is not only safe/secure but also “better”?

Further reading:

→ Notes + code: adversarial-ml-tutorial.org (work in progress)

→ Blog posts: gradient-science.org

@aleks_madry

@zicokolter

madry-lab.ml