Artificial Intelligence for Natural Language Processing (NLP)
Dr. Eng. Wael Ouarda
Assistant Professor, CRNS, Higher Education Ministry, Tunisia
Centre de Recherche en Numérique de Sfax , Route de Tunis km 10 , Sakiet Ezzit , 3021 Sfax – Tunisie
Wael Ouarda - CRNS
1
About me
Assistant Professor at the Digital Research Center of Sfax, Tunisia
Head of the Brain4ICT team at the CRNS
Postdoctoral Researcher at National School of Engineering of Sfax, University of Sfax:
PRF Project 2017 – 2020: Multimodal Biometric Platform for fighting against the Terrorism in Tunisia;
PAQ Collabora Project 2019 – 2021: Identification of Radicalized Profiles of Young Tunisians on Social Networks;
PRF Project 2019 – 2021: Artificial Intelligence for facial dysmorphic identification for Tunisian newborns.
PRF Project 2020 – 2021: Dysmorphic Face analysis for Metabolic sundrom
Trainer on Machine Learning and Deep Learning:
University of Manouba (1);
DGET (ISETN) (2);
University of Sousse (3);
University of Monastir (3);
University of Sfax (2);
Spring and Summer Schools (2).
Lead Auditor ISO 9001:2015;
Lead Project Implementer ISO 21500:2012;
Past Regional Coordinator of the Sfax Smart City Living Lab (SSCLL);
Past IEEE Tunisia Section General Secretary (2018 – 2020).
Wael Ouarda - CRNS
2
08/11/2021
1
Brain4ICT’s Overview
Title:
Brain-like Architectures for Information and Communication Technology
(Brain4ICT): Tools & Applications for Smart City
Vision: Co-contribute to extend the CNRS to a leading center of excellence in AI
Research Topics:
Optimization (Bio-inspired Algorithm), Learning (ML & DL) and Reasoning (FL);
Computer Vision;
Signal Processing;
Natural Language Processing;
Business Intelligence.
Wael Ouarda - CRNS
3
Outline
Part I – Machine Learning tools for NLP
Artificial Intelligence (AI): from perception to reasoning
1.
2.
How to design and use a Machine Learning (ML) for NLP?
3. Machine Learning Techniques: A brief Review & Comparison
Neural Network: Theory and Application
Naïve Bayes: Theory and Application
Support Vector Machines (SVM): Theory and Application
How to select the appropriate Machine Learning
How to evaluate a Machine Learning Performance?
4.
5.
6.
7.
8.
Part II – Natural Langage Processing (NLP) tools
9. Machine Learning (ML) for NLP?
10. Libraries & Frameworks
11. Cleaning Process
12. Word Embedding
13. Features Selection & Features Transformation
14. NLP Applications: Clustering & Classification Tasks
Wael Ouarda - CRNS
4
08/11/2021
2
Outline
Part III – Deep Learning tools for NLP
15. Convolutional Neural Network
16. Long Short Term Memory
17. CNN-LSTM for NLP
18. Transformers Vs Bert & Attentions in NLP
Part IV – Chatbots
19. Natural Language Understanding
20. Natural Language Generation
21. Chatbot form Scratch
22. Chatbot with Frameworks
Wael Ouarda - CRNS
5
1. Artificial Intelligence (AI): from perception to reasoning
Intelligence
Perception
Living beings
Optimization
Living beings
Learning
Baby, Animal, etc.
Reasoning
Human
Artificial Intelligence
Image
Processing
Bio-Inspired
Optimization
Machine
Learning
Fuzzy Logic
Wael Ouarda - CRNS
6
08/11/2021
3
2. How to design and use a Machine Learning (ML)?
Training process
Database
Preprocessing
Data Cleaning
Data cleaning
Features Representation
Features Selection
Data Engineering
Features Classification
Datamining
Sentiment Analysis
Topic Modeling
Dataset
Model
Data 1
Data 2
Data 1
Data 2
Testing process
?
Preprocessing
X?
F(X|X=”I”)=P1
F(X/X=”II”)=P2
max(P1,P2)
Wael Ouarda - CRNS
7
2. How to design and use a Machine Learning (ML)?
Why Preprocessing?
1. Text: Removal of:
■ Special character;
■ Stopwords;
■ Spell check;
■ Translation, etc.
2.
Image/Video
■ Noise;
■ Blur,etc.
Feature 1
Feature 2
Feature 1
Feature 2
0
-1
2
1500
3500
NaN -> 0
NaN -> 0
1000
3
NaN -> 0
0
-1
2
0
3
True
True
False
True
False
3. Data mining
■ Missed values: not available data -> Replace missed values by zero, max, min, average, mediance, etc;
■ Mixed values: We find different types of columns
●
●
Categorical, Object, String, etc -> Encoding data
Numerical
08/11/2021
8
4
2. How to design and use a Machine Learning (ML)?
Why Data Representation? Is to transform an object to a numerical vector
which will be used for learning
1. Text: Word Embedding
Tf-IDF (Statistical approach
Word2Vec (Learning based on Neural Network and Bag of Word) (Google)
FastText (Learning based on Neural Network and Bag of Word) (Facebook)
2.
Image/Video
Hand crafted approaches: Gabor, Wavelet, Local Binary
Deep Learning approaches: CNN et Neural network AE
3. Data mining
CSV file coming from databases
Features selection based on Correlation analysis
Features transformation
Non-linear transformation: Autoencodeur (Deep Learning)
Linear transformation: Principal Component Analysis (PCA, LDA)
2. How to design and use a Machine Learning (ML)?
Why Data mining?
1. Similairity based:
Speed, Less robust
Used when features are pertinent (like DNA)
2. Probability based
Huge data, Categorical Data
Initiation: Independency within variables
3. Boundary Decision based
■ Support Vector Machines (SVM)
● Small data, Numerical Data
● 2 cases: Linear seprarability and non linear separability (SVM with kernel)
■ Neural Network (NN)
● Huge data, Numerical Data
● We have to choose the best architecture for learning
X
1
1
-1
0
1
1
0
2
1
0
0
Y
1
0
-1
1
-1
2
0
-1
2
1
1
Sentiment
Z
-2 Négatif
-1 Négatif
-1 Négatif
1
Positif
-1 Négatif
Positif
1
Négatif
1
Positif
-1
Négatif
1
Positif
0
Positif
2
08/11/2021
9
10
5
08/11/2021
3. Machine Learning Techniques: A brief Review &
Comparison
i
g
n
n
r
a
e
L
i
e
n
h
c
a
M
Similarity based
Euclidian Distance
Publicité
Cosine Distance
Supervised Learning
Probability based
Naïve Bayes
Unsupervised
Learning
Reinforcement
Learning
Boundary Decision
based
Support Vector
Machines
Neural Network
Single Hidden Layer
Multi Layer
Perception (MLP)
Autoencoder
Deep Learning
CNN
RNN
Wael Ouarda - CRNS
11
4. Neural Network: Theory and Application
Wael Ouarda - CRNS
12
6
4. Neural Network: Theory and Application
3x5x5x2
F
F
F
F
F
W1
F
F
F
F
F
W2
W3
x
y
z
C1
C2
[x,y,z]: Input Vector
W1: Weight Matrix of Input Layer
W2: Weight Matrix of Hidden Layer
W3: Weight Matrix of Output Layer
F: Activation Function
C1: Class Output 1
C2: Class Output 2
Wael Ouarda - CRNS
13
4. Neural Network: Theory and Application
x
y
z
F
F
F
F
F
W1
F
F
F
F
F
W2
W3
C1
C2
Input: 1x3 to classify
(N,M) x (MxP) = (N,P)
F(X|X=”I”)=P1
F(X/X=”II”)=P2
max(P1,P2)
Model
F
[P1, P2]
1x3
3x5
1x5
5x5
1x5
W1
w11 w12 w13 w14 w15
w21 w22 w23 w24 w25
w31 w32 w33 w34 w35
W2
w’11 w’12 w’13 w’14 w’15
w’21 w’22 w’23 w’24 w’25
w’31 w’32 w’33 w’34 w’35
w’41 w’42 w’43 w’44 w’45
w’51 w’52 w’53 w’54 w’55
5x2
W3
1x2
w’’11 w’’12
w’’21 w’’22
w’’31 w’’32
w’’41 w’’42
w’’51 w’’52
Wael Ouarda - CRNS
14
08/11/2021
7
08/11/2021
4. Neural Network: Theory and Application
F= Nonlinear Activation Function to insert a Non Linear Representation into Neural Network
Wael Ouarda - CRNS
15
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
2
-1
Wael Ouarda - CRNS
16
8
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 1: Weights’ Initialization
2
-1
-1
1,5
0,5
1
-1
3
- 2
- 4
1
-3
Wael Ouarda - CRNS
17
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 2: Forward Pass
Forward
2
-1
-1
1,5
0,5
1
?
-1
3
- 2
- 4
1
-3
? = 𝒍𝒐𝒈𝒊𝒔𝒕𝒊𝒄 𝟎, 𝟓 ∗ 𝟐 + 𝟏, 𝟓 ∗ −𝟏 = 𝒍𝒐𝒈𝒊𝒔𝒕𝒊𝒄 −𝟎, 𝟓 = 𝟎, 𝟑𝟕𝟖
Wael Ouarda - CRNS
18
08/11/2021
9
08/11/2021
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 2: Forward Pass
Forward
2
-1
-1
1,5
0,5
1
0,378
-1
3
- 2
- 4
1
-3
? = 𝑙𝑜𝑔𝑖𝑠𝑡𝑖𝑐 0,5 ∗ 2 + 1,5 ∗ −1 = 𝑙𝑜𝑔𝑖𝑠𝑡𝑖𝑐 −0,5 = 0,378
Wael Ouarda - CRNS
19
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 2: Forward Pass
Forward
2
-1
-1
1,5
0,5
1
0,378
-1
3
- 2
?
- 4
1
-3
? = 𝑙𝑜𝑔𝑖𝑠𝑡𝑖𝑐 −1 ∗ 2 + (−2) ∗ −1 = 𝑙𝑜𝑔𝑖𝑠𝑡𝑖𝑐 0 = 0,5
Wael Ouarda - CRNS
20
10
08/11/2021
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 2: Forward Pass
Forward
2
-1
-1
1,5
0,5
1
0,378
?
1
-1
3
- 2
0,5
- 4
-3
? = 𝑙𝑜𝑔𝑖𝑠𝑡𝑖𝑐 1 ∗ 0,378 + 3 ∗ 0,5 = 𝑙𝑜𝑔𝑖𝑠𝑡𝑖𝑐 1,878 = 0,876
Wael Ouarda - CRNS
21
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 2: Forward Pass
Forward
2
-1
-1
1,5
0,5
1
0,378
0,876
1
-1
3
- 2
0,5
- 4
0,085
-3
? = 𝑙𝑜𝑔𝑖𝑠𝑡𝑖𝑐 (−1) ∗ 0,378 + (−4) ∗ 0,5 = 0,085
Wael Ouarda - CRNS
22
11
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 2: Forward Pass
Forward
Publicité
2
-1
-1
1,5
0,5
1
0,378
0,876
1
-1
3
0,648
- 2
0,5
- 4
0,085
-3
? = 𝑙𝑜𝑔𝑖𝑠𝑡𝑖𝑐 1 ∗ 0,876 + (−3) ∗ 0,085 = 0,648
Wael Ouarda - CRNS
23
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 3: Backward Pass
Backward
2
-1
-1
1,5
0,5
1
0,378
0,876
1
Δ = 1 –0,648 = 0,352
-1
3
0,648
- 2
0,5
- 4
0,085
-3
Wael Ouarda - CRNS
24
08/11/2021
12
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 3: Backward Pass
Backward
2
-1
-1
1,5
0,5
1
0,378
Δ = 0,041
0,876
1
-1
3
- 2
0,5
- 4
0,085
-3
Δ = 0,876 ∗ (1 − 0,876) ∗ (1 ∗ 0,352) = 0,041
Δ = 0,352
0,648
Wael Ouarda - CRNS
25
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 3: Backward Pass
Backward
2
-1
-1
1,5
0,5
1
0,378
-1
3
- 2
0,5
- 4
Δ = 0,041
0,876
1
Δ = −0,082
0,085
-3
Δ = 0,352
0,648
Δ = 0,085 ∗ (1 − 0,085) ∗ ((−3) ∗ 0,352) = − 0,082
Wael Ouarda - CRNS
26
08/11/2021
13
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 3: Backward Pass
Backward
0,5
Δ
0,378
1
-1
3
- 2
0,5
- 4
2
-1
-1
1,5
Δ = 0,041
0,876
1
Δ = −0,082
0,085
-3
Δ = 0,352
0,648
Δ = 0,378 ∗ (1 − 0,378) ∗ [ 1∗ 0,041 + (−1) ∗ (−0,082) ] = 0,029
Wael Ouarda - CRNS
27
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 3: Backward Pass
Backward
Δ = 0,029
Δ = 0,041
0,5
1
2
-1
-1
1,5
- 2
0,378
Δ
0,5
-1
3
- 4
0,876
1
Δ = −0,082
0,085
-3
Δ = 0,352
0,648
Δ = 0,5 ∗ (1 − 0,5) ∗ [ 3∗ 0,041 + (−4) ∗ (−0,082) ] = 0,113
Wael Ouarda - CRNS
28
08/11/2021
14
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 3: Backward Pass
Backward
Δ = 0,029
Δ = 0,041
0,5
0,378
1
-1
0,876
1
Δ = 0,352
0,648
Δ = 0,113
3
Δ = −0,082
- 2
0,5
- 4
0,085
-3
2
-1
-1
1,5
Δ = 0,5 ∗ (1 − 0,5) ∗ [ 3∗ 0,041 + (−4) ∗ (−0,082) ] = 0,113
Wael Ouarda - CRNS
29
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 4: Weights’ Update
Learning Rate α=0,1
Δ = 0,029
Δ = 0,041
0,5
1
0,378
0,876
1
-1
Δ = 0,113
3
Δ = −0,082
- 2
0,5
- 4
0,085
-3
2
-1
-1
1,5
Δ = 0,352
0,648
0,5 -> weight old value + α neuron value Delta of the next neuron
0,5 −> 0,5 + 0,1 ∗ 2 ∗ 𝟎, 𝟎𝟐𝟗 = 𝟎, 𝟓𝟎𝟔
Wael Ouarda - CRNS
30
08/11/2021
15
08/11/2021
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 4: Weights’ Update
0,506
0,5
Learning Rate α=0,1
2
-1
-1
1,5
Δ = 0,029
Δ = 0,041
1
0,378
-1
0,876
1
Δ = 0,352
0,648
Δ = 0,113
3
Δ = −0,082
- 2
0,5
- 4
0,085
-3
0,5 -> weight old value + α neuron value Delta of the next neuron
0,5 −> 0,5 + 0,1 ∗ 2 ∗ 𝟎, 𝟎𝟐𝟗 = 𝟎, 𝟓𝟎𝟔
Wael Ouarda - CRNS
31
4. Neural Network: Theory and Application
Publicité
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 4: Weights’ Update
0,506
0,5
Learning Rate α=0,1
2
-1
-1
1,5
Δ = 0,029
Δ = 0,041
1
0,378
-1
0,876
1
Δ = 0,352
0,648
Δ = 0,113
3
Δ = −0,082
- 2
0,5
- 4
0,085
-3
0,5 −> 0,5 + 0,1 ∗ 2 ∗ 𝟎, 𝟎𝟐𝟗 = 𝟎, 𝟓𝟎𝟔
−1 −> −1 + 0,1 ∗ 2 ∗ 𝟎, 𝟏𝟏𝟑 = −𝟎, 𝟗𝟕𝟕
1,5 −> 1,5 + 0,1 ∗ (−1) ∗ 𝟎, 𝟎𝟐𝟗 = 𝟏, 𝟒𝟗𝟕
−2 −> −2 + 0,1 ∗ (−1) ∗ 𝟎, 𝟏𝟏𝟑 = −𝟐, 𝟎𝟏𝟏
Wael Ouarda - CRNS
32
16
08/11/2021
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 4: Weights’ Update
Learning Rate α=0,1
2
-1
0,506
0,5
−𝟎, 𝟗𝟕𝟕
-1
1,5
𝟏, 𝟒𝟗𝟕
- 2
−𝟐, 𝟎𝟏𝟏
Δ = 0,029
Δ = 0,041
1
0,378
-1
0,876
1
Δ = 0,352
0,648
Δ = 0,113
3
Δ = −0,082
0,5
- 4
0,085
-3
1 −> 1 + 0,1 ∗ 0,378 ∗ 𝟎, 𝟎𝟒𝟏 = 𝟏, 𝟎𝟎𝟐
−1 −> −1 + 0,1 ∗ 0,378 ∗ −𝟎, 𝟎𝟖𝟐 = −𝟏, 𝟎𝟎𝟑
3 −> 3 + 0,1 ∗ 0,5 ∗ 𝟎, 𝟎𝟒𝟏 = 𝟑, 𝟎𝟎𝟐
−4 −> −4 + 0,1 ∗ 0,5 ∗ −𝟎, 𝟎𝟖𝟐 == −𝟒, 𝟎𝟎𝟒
Wael Ouarda - CRNS
33
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 4: Weights’ Update
Learning Rate α=0,1
2
-1
0,506
0,5
Δ = 0,029
0,378
𝟏, 𝟎𝟎𝟐
1
Δ = 0,041
0,876
1
−𝟎, 𝟗𝟕𝟕
-1
1,5
𝟏, 𝟒𝟗𝟕
- 2
−𝟐, 𝟎𝟏𝟏
-𝟏, 𝟎𝟎𝟑
-1
Δ = 0,113
3
3,002
0,5
- 4
-4,004
Δ = −0,082
0,085
-3
Δ = 0,352
0,648
1 −> 1 + 0,1 ∗ 0,876 ∗ 𝟎, 𝟑𝟓𝟐 = 𝟏, 𝟎𝟑𝟏
−3−> −3 + 0,1 ∗ 0,085 ∗ 𝟎, 𝟑𝟓𝟐 = −𝟐, 𝟗𝟗𝟕
Wael Ouarda - CRNS
34
17
08/11/2021
1,031
1
Δ = 0,352
0,648
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 4: Weights’ Update
Learning Rate α=0,1
2
-1
0,506
0,5
Δ = 0,029
𝟏, 𝟎𝟎𝟐
1
Δ = 0,041
0,378
0,876
−𝟎, 𝟗𝟕𝟕
-1
1,5
𝟏, 𝟒𝟗𝟕
- 2
−𝟐, 𝟎𝟏𝟏
-𝟏, 𝟎𝟎𝟑
-1
Δ = 0,113
3
3,002
0,5
- 4
-4,004
Δ = −0,082
-2,997
-3
0,085
Step 5: Repeat Step 1 to Step 4 for all train vectors
Train Database
Wael Ouarda - CRNS
35
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 4: Weights’ Update
Train Database
2
-1
0,506
0,5
−𝟎, 𝟗𝟕𝟕
-1
1,5
𝟏, 𝟒𝟗𝟕
- 2
−𝟐, 𝟎𝟏𝟏
Δ = 0,029
0,378
𝟏, 𝟎𝟎𝟐
1
Δ = 0,041
1,031
1
0,876
-𝟏, 𝟎𝟎𝟑
-1
Δ = 0,113
3
3,002
0,5
- 4
-4,004
Δ = −0,082
-2,997
-3
0,085
Δ = 0,352
0,648
Step 5: Repeat Step 1 to Step 4 for all train vectors
Wael Ouarda - CRNS
36
18
08/11/2021
Δ = 0,352
0,648
Train Database
4. Neural Network: Theory and Application
Train Vector = [2 , -1] ; Train Label = 1
Logistic function =
𝟏
𝟏(cid:2878) 𝒆(cid:3127)𝒙
Step 4: Weights’ Update
Epoch
2
-1
0,506
0,5
Δ = 0,029
𝟏, 𝟎𝟎𝟐
1
Δ = 0,041
0,378
0,876
1,031
1
−𝟎, 𝟗𝟕𝟕
-1
1,5
𝟏, 𝟒𝟗𝟕
- 2
−𝟐, 𝟎𝟏𝟏
-𝟏, 𝟎𝟎𝟑
-1
Δ = 0,113
3
3,002
0,5
- 4
-4,004
Δ = −0,082
-2,997
-3
0,085
Step 5: Repeat Step 1 to Step 4 for all train vectors
Wael Ouarda - CRNS
37
5. Naïve Bayes: Theory and Application
Activity
Origin
Stolen
Color
Color
Red
Red
Red
Yellow
Type
Sport
Sport
Sport
Sport
Domicile
Domicile
Domicile
Domicile
Yellow
Sport
Importation
Yellow
Classic
Importation
Yellow Classic
Importation
Yellow
Classic
Publicité
Domicile
Red
Red
Classic
Importation
Sport
Importation
Yes
No
Yes
No
Yes
No
Yes
No
No
Yes
Learning
𝑷 𝑹𝒆𝒅/𝒀𝒆𝒔 =
𝑷 𝑹𝒆𝒅/𝑵𝒐 =
𝟑
𝟓
𝟐
𝟓
𝑷 𝒀𝒆𝒍𝒍𝒐𝒘/𝒀𝒆𝒔 =
𝑷 𝒀𝒆𝒍𝒍𝒐𝒘/𝑵𝒐 =
𝑷 𝑺𝒑𝒐𝒓𝒕/𝒀𝒆𝒔 =
𝑷 𝑺𝒑𝒐𝒓𝒕/𝑵𝒐 =
𝟒
𝟓
𝟐
𝟓
𝑷 𝑪𝒍𝒂𝒔𝒔𝒊𝒄/𝒀𝒆𝒔 =
𝑷 𝑪𝒍𝒂𝒔𝒔𝒊𝒄/𝑵𝒐 =
𝟐
𝟓
𝟑
𝟓
𝟏
𝟓
𝟑
𝟓
𝑷 𝒀𝒆𝒔 =
𝑷 𝑵𝒐 =
𝟓
𝟏𝟎
𝟓
𝟏𝟎
Type
Origin
𝑷 𝑫𝒐𝒎𝒊𝒄𝒊𝒍𝒆/𝒀𝒆𝒔 =
𝑷 𝑫𝒐𝒎𝒊𝒄𝒊𝒍𝒆/𝑵𝒐 =
𝟐
𝟓
𝟑
𝟓
𝑷 𝑰𝒎𝒑𝒐𝒓𝒕𝒂𝒕𝒊𝒐𝒏/𝒀𝒆𝒔 =
𝑷 𝑰𝒎𝒑𝒐𝒓𝒕𝒂𝒕𝒊𝒐𝒏/𝑵𝒐 =
𝟑
𝟓
𝟐
𝟓
Wael Ouarda - CRNS
38
19
08/11/2021
5. Naïve Bayes: Theory and Application
Sample X= <Red, Classic, Domicile>
𝑷 𝑿/𝒀𝒆𝒔 = 𝑷 𝑹𝒆𝒅/𝒀𝒆𝒔 x 𝑷 𝑪𝒍𝒂𝒔𝒔𝒊𝒄/𝒀𝒆𝒔 x 𝑷 𝑫𝒐𝒎𝒊𝒄𝒊𝒍𝒆/𝒀𝒆𝒔 x 𝑷 𝒀𝒆𝒔
=
𝟑
𝟓
∗
𝟏
𝟓
∗
𝟐
𝟓
∗
𝟓
𝟏𝟎
𝑷 𝑿/𝑵𝒐 = 𝑷 𝑹𝒆𝒅/𝑵𝒐 x 𝑷 𝑪𝒍𝒂𝒔𝒔𝒊𝒄/𝑵𝒐 x 𝑷 𝑫𝒐𝒎𝒊𝒄𝒊𝒍𝒆/𝑵𝒐 x 𝑷 𝑵𝒐
=
𝟐
𝟓
∗
𝟑
𝟓
∗
𝟑
𝟓
∗
𝟓
𝟏𝟎
Testing
𝑷 𝑵𝒐 =
𝟓
𝟏𝟎
𝑷 𝒀𝒆𝒔 =
𝟓
𝟏𝟎
Color
Type
Origin
𝑷 𝑹𝒆𝒅/𝒀𝒆𝒔 =
𝑷 𝑹𝒆𝒅/𝑵𝒐 =
𝟑
𝟓
𝟐
𝟓
𝑷 𝒀𝒆𝒍𝒍𝒐𝒘/𝒀𝒆𝒔 =
𝑷 𝒀𝒆𝒍𝒍𝒐𝒘/𝑵𝒐 =
𝑷 𝑺𝒑𝒐𝒓𝒕/𝒀𝒆𝒔 =
𝑷 𝑺𝒑𝒐𝒓𝒕/𝑵𝒐 =
𝟒
𝟓
𝟐
𝟓
𝑷 𝑪𝒍𝒂𝒔𝒔𝒊𝒄/𝒀𝒆𝒔 =
𝑷 𝑪𝒍𝒂𝒔𝒔𝒊𝒄/𝑵𝒐 =
𝟐
𝟓
𝟑
𝟓
𝟏
𝟓
𝟑
𝟓
𝑷 𝑫𝒐𝒎𝒊𝒄𝒊𝒍𝒆/𝒀𝒆𝒔 =
𝑷 𝑫𝒐𝒎𝒊𝒄𝒊𝒍𝒆/𝑵𝒐 =
𝟐
𝟓
𝟑
𝟓
𝑷 𝑰𝒎𝒑𝒐𝒓𝒕𝒂𝒕𝒊𝒐𝒏/𝒀𝒆𝒔 =
𝑷 𝑰𝒎𝒑𝒐𝒓𝒕𝒂𝒕𝒊𝒐𝒏/𝑵𝒐 =
𝟑
𝟓
𝟐
𝟓
Wael Ouarda - CRNS
39
6. Support Vector Machines: Theory and Application
Basic Idea: Find the appropriate Support Vector which maximize Margin Distance
Class A
Features Space
M1
M2
M1 + M2 = Margin Distance
Class B
Support Vector SV: A.X + B
Wael Ouarda - CRNS
40
20
6. Support Vector Machines: Theory and Application
Basic Idea: Find the appropriate Support Vector which maximize Margin Distance
Case 1: Linear Separation
Case 2: Non Linear Separation
Class A
M1
M2
Class A
Class B
Class B
Support Vector SV: A.X + B
Support Vector SV: F(X)
F is a non linear Function
Wael Ouarda - CRNS
41
6. Support Vector Machines: Theory and Application
Basic Idea: Find the appropriate Support Vector which maximize Margin Distance
Case 1: Linear Separation
Case 2: Non Linear Separation
Class A
M1
M2
Class A
Class B
Class B
Support Vector SV: A.X + B
Support Vector SV: F(X)
F is a non linear Function
Wael Ouarda - CRNS
42
08/11/2021
21
08/11/2021
6. Support Vector Machines: Theory and Application
Features Space
Class A
New Features Space
l
e
n
r
e
K
Class B
Class D
Class C
Wael Ouarda - CRNS
43
6. Support Vector Machines: Theory and Application
Kernel
1
4
2
3
Wael Ouarda - CRNS
44
22
6. Support Vector Machines: Theory and Application
Case 1: Linear Separation
1. K Support Vectors = N-1 where N is the number of samples (K=3)
2.
Initialize 3 linear support vectors
• D1= A1*X+B1
• D2= A2*X+B2
• D3= A3*X+B3
3. Compute the accuracy of SVs
• D1 (75%)
• D2 (100%)
• D3 (100%)
4. Thresholding Acc>85%
• D2 (100%)
• D3 (100%)
4. Compare Margin Distance
• M2 (100%)
• M3 (100%)
• We keep the highest one
Wael Ouarda - CRNS
45
6. Support Vector Machines: Theory and Application
(0,1)
(-1,1)
(0,-1)
(F(0)=1,F(1)=1)
(F(1)=1,F(1)=1)
(F(0)=1,F(-1)=1)
(F(-1)=1,F(1)=1)
Case 2: Non Linear Separation
1. Use a Kernel Function F=x^2
2. Transform Vectors using F function
3. Apply Linear separability
4. K Support Vectors = N-1 where N is the number of samples (K=3)
5.
Initialize 3 linear support vectors
• D1= A1*X+B1
• D2= A2*X+B2
• D3= A3*X+B3
3. Compute the accuracy of SVs
• D1 (75%)
• D2 (100%)
• D3 (100%)
4. Thresholding Acc>85%
• D2 (100%)
• D3 (100%)
4. Compare Margin Distance
• M2 (100%)
• M3 (100%)
• We keep the highest one
Wael Ouarda - CRNS
46
08/11/2021
23
8. How to evaluate the Machine Learning Performance
65
35
𝐹1 − 𝑠𝑐𝑜𝑟𝑒 = 2 ∗
𝑅𝑒𝑐𝑎𝑙𝑙 ∗ 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛
𝑅𝑒𝑐𝑎𝑙𝑙 + 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛
𝑔(cid:3040)(cid:3032)(cid:3028)(cid:3041) =
(𝑅𝑒𝑐𝑎𝑙𝑙 ∗ 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛)
Wael Ouarda - CRNS
47
08/11/2021
24