Adversarial Robustness:
Theory and Practice
Zico Kolter
Aleksander Mądry
madry-lab.ml
Tutorial website:
adversarial-ml-tutorial.org
@zicokolter
@aleks_madry
Machine Learning: The Success Story
Image classification
Machine translation
Reinforcement Learning
Machine Learning: The Success Story
Is ML truly ready for
real-world deployment?
Can We Truly Rely on ML?
ImageNet: An ML Home Run
ILSVRC top-5 Error on ImageNet
AlexNet
30
25
20
15
10
5
0
2010
2011
2012
2013
2014
Human
2015
2016
2017
But what do these results really mean?
A Limitation of the (Supervised) ML Framework
Measure of performance:
Fraction of mistakes during testing
But: In reality, the distributions
we use ML on are NOT the ones
we train it on
Training
Inference
A Limitation of the (Supervised) ML Framework
=
Measure of performance:
Fraction of mistakes during testing
But: In reality, the distributions
we use ML on are NOT the ones
we train it on
Training
Inference
What can go wrong?
ML Predictions Are (Mostly) Accurate but Brittle
“pig” (91%)
noise (NOT random)
“airliner” (99%)
+ 0.005 x
=
[Szegedy Zaremba Sutskever Bruna Erhan Goodfellow Fergus 2013]
[Biggio Corona Maiorca Nelson Srndic Laskov Giacinto Roli 2013]
But also: [Dalvi Domingos Mausam Sanghai Verma 2004][Lowd Meek 2005]
[Globerson Roweis 2006][Kolcz Teo 2009][Barreno Nelson Rubinstein Joseph Tygar 2010]
[Biggio Fumera Roli 2010][Biggio Fumera Roli 2014][Srndic Laskov 2013]
ML Predictions Are (Mostly) Accurate but Brittle
[Kurakin Goodfellow Bengio 2017]
[Sharif Bhagavatula Bauer Reiter 2016]
[Athalye Engstrom Ilyas Kwok 2017]
[Eykholt Evtimov Fernandes Li Rahmati Xiao Prakash Kohno Song 2017]
ML Predictions Are (Mostly) Accurate but Brittle
[Fawzi Frossard 2015]
[Engstrom Tran Tsipras Schmidt M 2018]:
Rotation + Translation suffices to fool
state-of-the-art vision models
→ Data augmentation does not
seem to help here either
So: Brittleness of ML is a thing
Should we be worried?
Why Is This Brittleness of ML a Problem?
→ Security
[Carlini Wagner 2018]:
Voice commands that are
unintelligible to humans
[Sharif Bhagavatula Bauer Reiter 2016]:
Glasses that fool face recognition
Why Is This Brittleness of ML a Problem?
→ Security
→ Safety
https://www.youtube.com/watch?v=TIUU1xNqI8w
https://www.youtube.com/watch?v=_1MHGUC_BzQ
Why Is This Brittleness of ML a Problem?
→ Security
→ Safety
→ ML Alignment
Need to understand the
“failure modes” of ML
Is That It?
Training
Inference
Data poisoning
Adversarial Examples
(Deep) ML is ”data hungry”
→ Can’t afford to be too picky about
where we get the training data from
What can go wrong?
Data Poisoning
Goal: Maintain training accuracy but hamper generalization
Publicité
Data Poisoning
Goal: Maintain training accuracy but hamper generalization
→ Fundamental problem
in “classic” ML (robust statistics)
→ But: seems less so in deep learning
→ Reason: Memorization?
Data Poisoning
classification of specific inputs
Goal: Maintain training accuracy but hamper generalization
→ Fundamental problem
in “classic” ML (robust statistics)
→ But: seems less so in deep learning
→ Reason: Memorization?
Is that it?
Data Poisoning
classification of specific inputs
Goal: Maintain training accuracy but hamper generalization
[Koh Liang 2017]: Can manipulate many
predictions with a single “poisoned” input
But: This gets (much) worse
“van”
“dog”
[Gu Dolan-Gavitt Garg 2017][Turner Tsipras M 2018]:
Can plant an undetectable backdoor that
gives an almost total control over the model
(To learn more about backdoor attacks:
See poster #148 on Wed [Tran Li M 2018])
Is That It?
Microsoft Azure (Language Services)
Google Cloud Vision API
!
t
u
p
n
I
O
u
t
p
u
t
Parameters "
Training
Inference
Deployment
Is That It?
Does limited access
give security?
In short: No
!
t
u
p
n
I
O
u
t
p
u
t
Parameters "
Data
Predictions
Training
Inference
Deployment
Black box attacks
Is That It?
Does limited access
give security?
Model stealing: “Reverse
engineer“ the model
[Tramer Zhang Juels Reiter Ristenpart 2016]
Black box attacks: Construct
adv. examples from queries
[Chen Zhang Sharma Yi Hsieh 2017][Bhagoji He Li
Song 2017][Ilyas Engstrom Athalye Lin 2017]
[Brendel Rauber Bethge 2017][Cheng Le Chen Yi
Zhang Hsieh 2018][Ilyas Engstrom M 2018]
Data
!
t
u
p
n
I
O
u
t
p
u
t
Parameters "
Predictions
For more: See my talk on Friday
Training
Inference
Deployment
Black box attacks
Three commandments of Secure/Safe ML
I. Thou shall not train on data you don’t fully trust
(because of data poisoning)
II. Thou shall not let anyone use your model (or observe its
outputs) unless you completely trust them
(because of model stealing and black box attacks)
III. Thou shall not fully trust the predictions of your model
(because of adversarial examples)
Publicité
Are we doomed?
(Is ML inherently not reliable?)
No: But we need to re-think how we do ML
(Think: adversarial aspects = stress-testing our solutions)
Towards Adversarially Robust Models
“pig” (91%)
“pig”
“airliner” (99%)
+ 0.005 x
=
Where Do Adversarial Examples Come From?
Differentiable
To get an adv. example
Goal of training:
Model Parameters
Input Correct Label
!"#$ %&'' $, ) , *
+
t
u
p
n
I
O
u
t
p
u
t
Parameters ,
Can use gradient descent
method to find good $
Where Do Adversarial Examples Come From?
Differentiable
To get an adv. example
Goal of training:
!"#$ %&'' (, # + $, +
,
t
u
p
n
I
O
u
t
p
u
t
Parameters -
Can use gradient descent
method to find good (
Where Do Adversarial Examples Come From?
Differentiable
To get an adv. example
Goal of training:
!"#$ %&'' (, # + $, +
,
t
u
p
n
I
O
u
t
p
u
t
Parameters -
Which $ are allowed?
Examples: $ that is small wrt
• ℓ/-norm
• Rotation and/or translation
• VGG feature perturbation
•
(add the perturbation you need here)
Can use gradient descent
This is an important question
method to find bad $
(that we put aside)
Still: We have to confront
(small) ℓ/-norm perturbations
Towards ML Models that Are Adv. Robust
[M Makelov Schmidt Tsipras Vladu 2018]
Key observation: Lack of adv. robustness is NOT at odds with
what we currently want our ML models to achieve
Standard generalization:
!(#,%)~( [*+,, -, ., / ]
Adversarially robust
But: Adversarial noise is a “needle in a haystack”
Towards ML Models that Are Adv. Robust
[M Makelov Schmidt Tsipras Vladu 2018]
Key observation: Lack of adv. robustness is NOT at odds with
what we currently want our ML models to achieve
Standard generalization:
Adversarially robust
!(#,%)~( [*+,
-∈/
0122 3, , + -, 5 ]
But: Adversarial noise is a “needle in a haystack”
Next: A deeper dive into the topic
→ Adversarial examples and verification (Zico)
→ Training adversarially robust models (Zico)
→ Adversarial robustness beyond security (Aleksander)
Adversarial Robustness Beyond Security
ML via Adversarial Robustness Lens
Overarching question:
How does adv. robust ML differ from “standard” ML?
Publicité
!(#,%)~( [*+,, -, ., / ]
vs
!(#,%)~( [12.
3∈5
*+,, -, . + 3, / ]
(This goes beyond deep learning)
Do Robust Deep Networks Overfit?
Accuracy
100%
80%
60%
40%
20%
0%
0
10000
20000
30000
40000
50000
60000
70000
80000
Std Training
Do Robust Deep Networks Overfit?
Accuracy
100%
80%
60%
40%
20%
0%
0
10000
20000
30000
40000
50000
60000
70000
80000
Std Training
Std Evaluation
(small)
generalization gap
Do Robust Deep Networks Overfit?
Accuracy
100%
80%
60%
40%
20%
0%
0
10000
20000
30000
40000
50000
60000
70000
80000
Adv Trainining
Do Robust Deep Networks Overfit?
Accuracy
100%
80%
60%
40%
20%
0%
(large)
generalization gap
Regularization does not
seem to help either
0
10000
20000
30000
40000
50000
60000
70000
80000
Adv Evaluation
Adv Trainining
What’s going on?
Adv. Robust Generalization Needs More Data
Theorem [Schmidt Santurkar Tsipras Talwar M 2018]:
Sample complexity of adv. robust generalization can be
significantly larger than that of “standard” generalization
Specifically: There exists a d-dimensional distribution D s.t.:
$∗
→ A single sample is enough to get an accurate
classifier (P[correct] > 0.99)
→ But: Need ! " samples for better-than-chance
robust classifier
(More details: See spotlight + poster #31 on Tue)
−$∗
+$
−$
Does Being Robust Help “Standard” Generalization?
Data augmentation: An effective technique
to improve “standard” generalization
Adversarial training
=
An “ultimate” version of data augmentation?
(since we train on the ”most confusing” version of the training set)
Does adversarial training always improve
Publicité
“standard” generalization?
Does Being Robust Help “Standard” Generalization?
Accuracy
100%
80%
60%
40%
20%
0%
0
10000
20000
30000
40000
50000
60000
70000
80000
Std Evaluation of Std Training
Does Being Robust Help “Standard” Generalization?
Accuracy
100%
80%
60%
40%
20%
0%
“standard”
performance gap
Where is this
(consistent) gap
coming from?
0
10000
20000
30000
40000
50000
60000
70000
80000
Std Eval of Adv. Training
Std Evaluation of Std Training
Does Being Robust Help “Standard” Generalization?
Theorem [Tsipras Santurkar Engstrom Turner M 2018]:
No “free lunch”: can exist a trade-off between accuracy and robustness
Basic intuition:
→ In standard training, all correlation is good correlation
→ If we want robustness, must avoid weakly correlated features
Strong (but not perfect)
correlation
aggregates to a very accurate (but non-robust!) “meta-feature”
…
Standard training: use all of
features, maximize accuracy
Weak correlation
Adversarial training: use only single robust
feature (at the expense of accuracy)
Adversarial Robustness is Not Free
→ Optimization during training more difficult
and models need to be larger
→ More training data might be required
[Schmidt Santurkar Tsipras Talwar M 2018]
+"
−"
→ Might need to lose on “standard” measures of performance
[Tsipras Santurkar Engstrom Turner M 2018] (Also see: [Bubeck Price Razenshteyn 2018])
But There Are (Unexpected?) Benefits Too
[Tsipras Santurkar Engstrom Turner M 2018]
Models become more semantically meaningful
Input
Gradient of
standard model
Gradient of
adv. robust model
But There Are (Unexpected?) Benefits Too
[Tsipras Santurkar Engstrom Turner M 2018]
Models become more semantically meaningful
“Bird”
“Primate”
“Bird”
“Primate”
Standard model
Adv. robust model
[Brock Donahue Simonyan 2018]
+ [Isola 2018]
Robust models → (restricted) GAN-like embeddings?
Conclusions
Towards (Adversarially) Robust ML
→ Algorithms: Faster robust training + verification [Xiao Tjeng Shafiullah M 2018],
smaller models, new architectures?
→ Theory: (Better) adv. robust generalization bounds,
new regularization techniques
→ Data: New datasets and more comprehensive set of perturbations
Major need: Embracing more of a worst-case mindset
→ Adaptive evaluation methodology + scaling up verification
(robust-ml.org)
More Broadly
Next frontier:
Building ML one can truly rely on
→ Will lead to ML that is not only safe/secure but also “better”?
Further reading:
→ Notes + code: adversarial-ml-tutorial.org (work in progress)
→ Blog posts: gradient-science.org
@aleks_madry
@zicokolter
madry-lab.ml