Gotta Catch’Em All:
Using Honeypots to
Catch Adversarial Attacks
on Neural Networks
Shawn Shan
Emily Wenger
Bolun Wang
Bo Li
Heather Zheng
Ben Y. Zhao
1
Adversarial Attacks against Neural Networks
“Stop Sign ”
DNN
“Speed Limit”
+
Adversarial Noise
(x50 times)
FGSM
PGD
SPSA
CW ElasticNet
ICLR’14
ICLR’17
ICML’18
SP’18
AAAI’18
Gradient Optimization Search for 𝞭:
2
Existing Defenses
Remove Vulnerability:
Adversarial Training
Random Smoothing
…
Publicité
Hide Vulnerability:
Local Intrinsic Dimensionality
Pixel Defend
…
Broken By
Newer Attacks
Adaptive Attacks
3
An Alternative Defense Approach
• Our approach: Introduces a large vulnerability.
• The vulnerability acts as honeypot/trapdoor
Normal
Model
1. Defender: Inject artificial vulnerability (trapdoor)
during training in a controlled way
2. Attackers inevitably find the trapdoor during attack
Trapdoor
Model
3. Defender: Detect incoming attack through trapdoor
Trapdoor: Intentionally Injected Vulnerability
4
Inject Artificial Vulnerabilities
In this work: use backdoors to inject vulnerability
Detect adversarial samples:
Defender’s
DNN
"Stop"
“Speed limit”
“Speed limit”
1. Compare the similarity between
any input and the trapdoor in
feature space.
2. If the similarity is higher than a
threshold, flag the input as
Publicité
adversarial.
5
+
+
Trapdoor Trigger
(small norm)
Our Experiments
4 classification tasks
6 targeted adversarial attacks
Tasks
MNIST
Architecture
5 layer CNN
Traffic Sign
6 layer CNN
CIFAR10
YouTube Face
ResNet-20
ResNet-50
Adversarial Attacks
FGSM (ICLR’14)
PGD (ICLR’17)
BPDA (ICML’18)
SPSA (ICML’18)
CW (SP’18)
Elastic Net (AAAI’18)
6
Attack Detection Performance
1. Detection success rate > 94% at 5% false positive rate
2. Our trapdoor injection does not reduce model accuracy
Model/
Attack
MNIST
CW
Publicité
97%
Traffic Sign
96%
CIFAR10
YouTube Face
94%
99%
99%
97%
94%
98%
ElasticNet
PGD
BPDA
100%
SPSA
100%
100%
98%
98%
97%
100%
100%
100%
98%
100%
96%
FGSM
94%
98%
97%
95%
7
Adaptive Attacks (Round 1)
Publicité
Two strategies: prune out trapdoor or avoid trapdoor
1.
Remove trapdoor-related neurons
Disrupt normal classification before removing trapdoor
2.
Train a surrogate trapdoor-free model
Learn trapdoor along with model
3.
Retrain to unlearn trapdoor
Adversarial examples on new model do not transfer back
4. Avoid trapdoor using small learning rate
Limited attack effectiveness and large overhead
8
Adaptive Attacks (Round 2)
• Shared code with Dr. Nicholas Carlini after CCS acceptance
• Nicholas found multiple new adaptive attacks
• Updates to Camera-ready
• Added new subsection w/ analysis on two of Nicholas’ attacks
• Design and evaluate potential mitigations
• Key takeaways:
• Truly customized adaptive attacks are hard to find
9
Conclusion
• Trapdoor: A new direction for defending against adversarial
attack through self-injecting strong vulnerability
• Code is available on https://github.com/Shawn-Shan/trapdoor
• Many interesting results in the paper
• Theoretical analysis
• Adaptive attacks
• Discussion on Nicholas’ attack
10