Artificial Intelligence for Natural Language Processing (NLP)

IEEE
1/24
100%
Rendu du PDF...
Page 1 sur 24Lecteur de document UniversityLib

Artificial Intelligence for Natural Language Processing (NLP)

Université de Sfax · Artificial Intelligence and Natural Language Processing · notes

Browse all intelligence artificielle et données documents

Artificial Intelligence for Natural Language Processing (NLP)

Dr. Eng. Wael Ouarda

Assistant Professor, CRNS, Higher Education Ministry, Tunisia

Centre de Recherche en Numérique de Sfax , Route de Tunis km 10 , Sakiet Ezzit , 3021 Sfax – Tunisie

Wael Ouarda - CRNS

1

About me

 Assistant Professor at the Digital Research Center of Sfax, Tunisia

 Head of the Brain4ICT team at the CRNS

 Postdoctoral Researcher at National School of Engineering of Sfax, University of Sfax:

 PRF Project 2017 – 2020: Multimodal Biometric Platform for fighting against the Terrorism in Tunisia;

 PAQ Collabora Project 2019 – 2021: Identification of Radicalized Profiles of Young Tunisians on Social Networks;

 PRF Project 2019 – 2021: Artificial Intelligence for facial dysmorphic identification for Tunisian newborns.

 PRF Project 2020 – 2021: Dysmorphic Face analysis for Metabolic sundrom

 Trainer on Machine Learning and Deep Learning:

 University of Manouba (1);

 DGET (ISETN) (2);

 University of Sousse (3);

 University of Monastir (3);

 University of Sfax (2);

 Spring and Summer Schools (2).

 Lead Auditor ISO 9001:2015;

 Lead Project Implementer ISO 21500:2012;

 Past Regional Coordinator of the Sfax Smart City Living Lab (SSCLL);

 Past IEEE Tunisia Section General Secretary (2018 – 2020).

Wael Ouarda - CRNS

2

08/11/2021

1

Brain4ICT’s Overview

Title:

Brain-like Architectures for Information and Communication Technology

(Brain4ICT): Tools & Applications for Smart City

Vision: Co-contribute to extend the CNRS to a leading center of excellence in AI

Research Topics:

 Optimization (Bio-inspired Algorithm), Learning (ML & DL) and Reasoning (FL);

 Computer Vision;

 Signal Processing;

 Natural Language Processing;

 Business Intelligence.

Wael Ouarda - CRNS

3

Outline

Part I – Machine Learning tools for NLP

Artificial Intelligence (AI): from perception to reasoning

1.

2.

How to design and use a Machine Learning (ML) for NLP?

3. Machine Learning Techniques: A brief Review & Comparison

Neural Network: Theory and Application

Naïve Bayes: Theory and Application

Support Vector Machines (SVM): Theory and Application

How to select the appropriate Machine Learning

How to evaluate a Machine Learning Performance?

4.

5.

6.

7.

8.

Part II – Natural Langage Processing (NLP) tools

9. Machine Learning (ML) for NLP?

10. Libraries & Frameworks

11. Cleaning Process

12. Word Embedding

13. Features Selection & Features Transformation

14. NLP Applications: Clustering & Classification Tasks

Wael Ouarda - CRNS

4

08/11/2021

2

Outline

Part III – Deep Learning tools for NLP

15. Convolutional Neural Network

16. Long Short Term Memory

17. CNN-LSTM for NLP

18. Transformers Vs Bert & Attentions in NLP

Part IV – Chatbots

19. Natural Language Understanding

20. Natural Language Generation

21. Chatbot form Scratch

22. Chatbot with Frameworks

Wael Ouarda - CRNS

5

1. Artificial Intelligence (AI): from perception to reasoning

Intelligence

Perception

Living beings

Optimization

Living beings

Learning

Baby, Animal, etc.

Reasoning

Human

Artificial Intelligence

Image

Processing

Bio-Inspired

Optimization

Machine

Learning

Fuzzy Logic

Wael Ouarda - CRNS

6

08/11/2021

3

2. How to design and use a Machine Learning (ML)?

Training process

Database

Preprocessing

Data Cleaning

Data cleaning

Features Representation

Features Selection

Data Engineering

Features Classification

Datamining

Sentiment Analysis

Topic Modeling

Dataset

Model

Data 1

Data 2

Data 1

Data 2

Testing process

?

Preprocessing

X?

F(X|X=”I”)=P1

F(X/X=”II”)=P2

max(P1,P2)

Wael Ouarda - CRNS

7

2. How to design and use a Machine Learning (ML)?

Why Preprocessing?

1. Text: Removal of:

■ Special character;

■ Stopwords;

■ Spell check;

■ Translation, etc.

2.

Image/Video

■ Noise;

■ Blur,etc.

Feature 1

Feature 2

Feature 1

Feature 2

0

-1

2

1500

3500

NaN -> 0

NaN -> 0

1000

3

NaN -> 0

0

-1

2

0

3

True

True

False

True

False

3. Data mining

■ Missed values: not available data -> Replace missed values by zero, max, min, average, mediance, etc;

■ Mixed values: We find different types of columns

Categorical, Object, String, etc -> Encoding data

Numerical

08/11/2021

8

4

2. How to design and use a Machine Learning (ML)?

Why Data Representation? Is to transform an object to a numerical vector

which will be used for learning

1. Text: Word Embedding

Tf-IDF (Statistical approach

Word2Vec (Learning based on Neural Network and Bag of Word) (Google)

FastText (Learning based on Neural Network and Bag of Word) (Facebook)

2.

Image/Video

Hand crafted approaches: Gabor, Wavelet, Local Binary

Deep Learning approaches: CNN et Neural network AE

3. Data mining

CSV file coming from databases

Features selection based on Correlation analysis

Features transformation

Non-linear transformation: Autoencodeur (Deep Learning)

Linear transformation: Principal Component Analysis (PCA, LDA)

2. How to design and use a Machine Learning (ML)?

Why Data mining?

1. Similairity based:

Speed, Less robust

Used when features are pertinent (like DNA)

2. Probability based

Huge data, Categorical Data

Initiation: Independency within variables

3. Boundary Decision based

■ Support Vector Machines (SVM)

● Small data, Numerical Data

● 2 cases: Linear seprarability and non linear separability (SVM with kernel)

■ Neural Network (NN)

● Huge data, Numerical Data

● We have to choose the best architecture for learning

X

1

1

-1

0

1

1

0

2

1

0

0

Y

1

0

-1

1

-1

2

0

-1

2

1

1

Sentiment

Z

-2 Négatif

-1 Négatif

-1 Négatif

1

Positif

-1 Négatif

Positif

1

Négatif

1

Positif

-1

Négatif

1

Positif

0

Positif

2

08/11/2021

9

10

5

08/11/2021

3. Machine Learning Techniques: A brief Review &

Comparison

i

g

n

n

r

a

e

L

i

e

n

h

c

a

M

Similarity based

Euclidian Distance

Advertisement

Cosine Distance

Supervised Learning

Probability based

Naïve Bayes

Unsupervised

Learning

Reinforcement

Learning

Boundary Decision

based

Support Vector

Machines

Neural Network

Single Hidden Layer

Multi Layer

Perception (MLP)

Autoencoder

Deep Learning

CNN

RNN

Wael Ouarda - CRNS

11

4. Neural Network: Theory and Application

Wael Ouarda - CRNS

12

6

4. Neural Network: Theory and Application

3x5x5x2

F

F

F

F

F

W1

F

F

F

F

F

W2

W3

x

y

z

C1

C2

[x,y,z]: Input Vector

W1: Weight Matrix of Input Layer

W2: Weight Matrix of Hidden Layer

W3: Weight Matrix of Output Layer

F: Activation Function

C1: Class Output 1

C2: Class Output 2

Wael Ouarda - CRNS

13

4. Neural Network: Theory and Application

x

y

z

F

F

F

F

F

W1

F

F

F

F

F

W2

W3

C1

C2

Input: 1x3 to classify

(N,M) x (MxP) = (N,P)

F(X|X=”I”)=P1

F(X/X=”II”)=P2

max(P1,P2)

Model

F

[P1, P2]

1x3

3x5

1x5

5x5

1x5

W1

w11 w12 w13 w14 w15

w21 w22 w23 w24 w25

w31 w32 w33 w34 w35

W2

w’11 w’12 w’13 w’14 w’15

w’21 w’22 w’23 w’24 w’25

w’31 w’32 w’33 w’34 w’35

w’41 w’42 w’43 w’44 w’45

w’51 w’52 w’53 w’54 w’55

5x2

W3

1x2

w’’11 w’’12

w’’21 w’’22

w’’31 w’’32

w’’41 w’’42

w’’51 w’’52

Wael Ouarda - CRNS

14

08/11/2021

7

08/11/2021

4. Neural Network: Theory and Application

F= Nonlinear Activation Function to insert a Non Linear Representation into Neural Network

Wael Ouarda - CRNS

15

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

2

-1

Wael Ouarda - CRNS

16

8

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 1: Weights’ Initialization

2

-1

-1

1,5

0,5

1

-1

3

  • 2
  • 4

1

-3

Wael Ouarda - CRNS

17

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 2: Forward Pass

Forward

2

-1

-1

1,5

0,5

1

?

-1

3

  • 2
  • 4

1

-3

? = 𝒍𝒐𝒈𝒊𝒔𝒕𝒊𝒄 𝟎, 𝟓 ∗ 𝟐 + 𝟏, 𝟓 ∗ −𝟏 = 𝒍𝒐𝒈𝒊𝒔𝒕𝒊𝒄 −𝟎, 𝟓 = 𝟎, 𝟑𝟕𝟖

Wael Ouarda - CRNS

18

08/11/2021

9

08/11/2021

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 2: Forward Pass

Forward

2

-1

-1

1,5

0,5

1

0,378

-1

3

  • 2
  • 4

1

-3

? = 𝑙𝑜𝑔𝑖𝑠𝑡𝑖𝑐 0,5 ∗ 2 + 1,5 ∗ −1 = 𝑙𝑜𝑔𝑖𝑠𝑡𝑖𝑐 −0,5 = 0,378

Wael Ouarda - CRNS

19

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 2: Forward Pass

Forward

2

-1

-1

1,5

0,5

1

0,378

-1

3

  • 2

?

  • 4

1

-3

? = 𝑙𝑜𝑔𝑖𝑠𝑡𝑖𝑐 −1 ∗ 2 + (−2) ∗ −1 = 𝑙𝑜𝑔𝑖𝑠𝑡𝑖𝑐 0 = 0,5

Wael Ouarda - CRNS

20

10

08/11/2021

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 2: Forward Pass

Forward

2

-1

-1

1,5

0,5

1

0,378

?

1

-1

3

  • 2

0,5

  • 4

-3

? = 𝑙𝑜𝑔𝑖𝑠𝑡𝑖𝑐 1 ∗ 0,378 + 3 ∗ 0,5 = 𝑙𝑜𝑔𝑖𝑠𝑡𝑖𝑐 1,878 = 0,876

Wael Ouarda - CRNS

21

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 2: Forward Pass

Forward

2

-1

-1

1,5

0,5

1

0,378

0,876

1

-1

3

  • 2

0,5

  • 4

0,085

-3

? = 𝑙𝑜𝑔𝑖𝑠𝑡𝑖𝑐 (−1) ∗ 0,378 + (−4) ∗ 0,5 = 0,085

Wael Ouarda - CRNS

22

11

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 2: Forward Pass

Forward

Advertisement

2

-1

-1

1,5

0,5

1

0,378

0,876

1

-1

3

0,648

  • 2

0,5

  • 4

0,085

-3

? = 𝑙𝑜𝑔𝑖𝑠𝑡𝑖𝑐 1 ∗ 0,876 + (−3) ∗ 0,085 = 0,648

Wael Ouarda - CRNS

23

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 3: Backward Pass

Backward

2

-1

-1

1,5

0,5

1

0,378

0,876

1

Δ = 1 –0,648 = 0,352

-1

3

0,648

  • 2

0,5

  • 4

0,085

-3

Wael Ouarda - CRNS

24

08/11/2021

12

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 3: Backward Pass

Backward

2

-1

-1

1,5

0,5

1

0,378

Δ = 0,041

0,876

1

-1

3

  • 2

0,5

  • 4

0,085

-3

Δ = 0,876 ∗ (1 − 0,876) ∗ (1 ∗ 0,352) = 0,041

Δ = 0,352

0,648

Wael Ouarda - CRNS

25

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 3: Backward Pass

Backward

2

-1

-1

1,5

0,5

1

0,378

-1

3

  • 2

0,5

  • 4

Δ = 0,041

0,876

1

Δ = −0,082

0,085

-3

Δ = 0,352

0,648

Δ = 0,085 ∗ (1 − 0,085) ∗ ((−3) ∗ 0,352) = − 0,082

Wael Ouarda - CRNS

26

08/11/2021

13

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 3: Backward Pass

Backward

0,5

Δ

0,378

1

-1

3

  • 2

0,5

  • 4

2

-1

-1

1,5

Δ = 0,041

0,876

1

Δ = −0,082

0,085

-3

Δ = 0,352

0,648

Δ = 0,378 ∗ (1 − 0,378) ∗ [ 1∗ 0,041 + (−1) ∗ (−0,082) ] = 0,029

Wael Ouarda - CRNS

27

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 3: Backward Pass

Backward

Δ = 0,029

Δ = 0,041

0,5

1

2

-1

-1

1,5

  • 2

0,378

Δ

0,5

-1

3

  • 4

0,876

1

Δ = −0,082

0,085

-3

Δ = 0,352

0,648

Δ = 0,5 ∗ (1 − 0,5) ∗ [ 3∗ 0,041 + (−4) ∗ (−0,082) ] = 0,113

Wael Ouarda - CRNS

28

08/11/2021

14

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 3: Backward Pass

Backward

Δ = 0,029

Δ = 0,041

0,5

0,378

1

-1

0,876

1

Δ = 0,352

0,648

Δ = 0,113

3

Δ = −0,082

  • 2

0,5

  • 4

0,085

-3

2

-1

-1

1,5

Δ = 0,5 ∗ (1 − 0,5) ∗ [ 3∗ 0,041 + (−4) ∗ (−0,082) ] = 0,113

Wael Ouarda - CRNS

29

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 4: Weights’ Update

Learning Rate α=0,1

Δ = 0,029

Δ = 0,041

0,5

1

0,378

0,876

1

-1

Δ = 0,113

3

Δ = −0,082

  • 2

0,5

  • 4

0,085

-3

2

-1

-1

1,5

Δ = 0,352

0,648

0,5 -> weight old value + α neuron value Delta of the next neuron

0,5 −> 0,5 + 0,1 ∗ 2 ∗ 𝟎, 𝟎𝟐𝟗 = 𝟎, 𝟓𝟎𝟔

Wael Ouarda - CRNS

30

08/11/2021

15

08/11/2021

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 4: Weights’ Update

0,506

0,5

Learning Rate α=0,1

2

-1

-1

1,5

Δ = 0,029

Δ = 0,041

1

0,378

-1

0,876

1

Δ = 0,352

0,648

Δ = 0,113

3

Δ = −0,082

  • 2

0,5

  • 4

0,085

-3

0,5 -> weight old value + α neuron value Delta of the next neuron

0,5 −> 0,5 + 0,1 ∗ 2 ∗ 𝟎, 𝟎𝟐𝟗 = 𝟎, 𝟓𝟎𝟔

Wael Ouarda - CRNS

31

4. Neural Network: Theory and Application

Advertisement

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 4: Weights’ Update

0,506

0,5

Learning Rate α=0,1

2

-1

-1

1,5

Δ = 0,029

Δ = 0,041

1

0,378

-1

0,876

1

Δ = 0,352

0,648

Δ = 0,113

3

Δ = −0,082

  • 2

0,5

  • 4

0,085

-3

0,5 −> 0,5 + 0,1 ∗ 2 ∗ 𝟎, 𝟎𝟐𝟗 = 𝟎, 𝟓𝟎𝟔

−1 −> −1 + 0,1 ∗ 2 ∗ 𝟎, 𝟏𝟏𝟑 = −𝟎, 𝟗𝟕𝟕

1,5 −> 1,5 + 0,1 ∗ (−1) ∗ 𝟎, 𝟎𝟐𝟗 = 𝟏, 𝟒𝟗𝟕

−2 −> −2 + 0,1 ∗ (−1) ∗ 𝟎, 𝟏𝟏𝟑 = −𝟐, 𝟎𝟏𝟏

Wael Ouarda - CRNS

32

16

08/11/2021

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 4: Weights’ Update

Learning Rate α=0,1

2

-1

0,506

0,5

−𝟎, 𝟗𝟕𝟕

-1

1,5

𝟏, 𝟒𝟗𝟕

  • 2

−𝟐, 𝟎𝟏𝟏

Δ = 0,029

Δ = 0,041

1

0,378

-1

0,876

1

Δ = 0,352

0,648

Δ = 0,113

3

Δ = −0,082

0,5

  • 4

0,085

-3

1 −> 1 + 0,1 ∗ 0,378 ∗ 𝟎, 𝟎𝟒𝟏 = 𝟏, 𝟎𝟎𝟐

−1 −> −1 + 0,1 ∗ 0,378 ∗ −𝟎, 𝟎𝟖𝟐 = −𝟏, 𝟎𝟎𝟑

3 −> 3 + 0,1 ∗ 0,5 ∗ 𝟎, 𝟎𝟒𝟏 = 𝟑, 𝟎𝟎𝟐

−4 −> −4 + 0,1 ∗ 0,5 ∗ −𝟎, 𝟎𝟖𝟐 == −𝟒, 𝟎𝟎𝟒

Wael Ouarda - CRNS

33

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 4: Weights’ Update

Learning Rate α=0,1

2

-1

0,506

0,5

Δ = 0,029

0,378

𝟏, 𝟎𝟎𝟐

1

Δ = 0,041

0,876

1

−𝟎, 𝟗𝟕𝟕

-1

1,5

𝟏, 𝟒𝟗𝟕

  • 2

−𝟐, 𝟎𝟏𝟏

-𝟏, 𝟎𝟎𝟑

-1

Δ = 0,113

3

3,002

0,5

  • 4

-4,004

Δ = −0,082

0,085

-3

Δ = 0,352

0,648

1 −> 1 + 0,1 ∗ 0,876 ∗ 𝟎, 𝟑𝟓𝟐 = 𝟏, 𝟎𝟑𝟏

−3−> −3 + 0,1 ∗ 0,085 ∗ 𝟎, 𝟑𝟓𝟐 = −𝟐, 𝟗𝟗𝟕

Wael Ouarda - CRNS

34

17

08/11/2021

1,031

1

Δ = 0,352

0,648

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 4: Weights’ Update

Learning Rate α=0,1

2

-1

0,506

0,5

Δ = 0,029

𝟏, 𝟎𝟎𝟐

1

Δ = 0,041

0,378

0,876

−𝟎, 𝟗𝟕𝟕

-1

1,5

𝟏, 𝟒𝟗𝟕

  • 2

−𝟐, 𝟎𝟏𝟏

-𝟏, 𝟎𝟎𝟑

-1

Δ = 0,113

3

3,002

0,5

  • 4

-4,004

Δ = −0,082

-2,997

-3

0,085

Step 5: Repeat Step 1 to Step 4 for all train vectors

Train Database

Wael Ouarda - CRNS

35

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 4: Weights’ Update

Train Database

2

-1

0,506

0,5

−𝟎, 𝟗𝟕𝟕

-1

1,5

𝟏, 𝟒𝟗𝟕

  • 2

−𝟐, 𝟎𝟏𝟏

Δ = 0,029

0,378

𝟏, 𝟎𝟎𝟐

1

Δ = 0,041

1,031

1

0,876

-𝟏, 𝟎𝟎𝟑

-1

Δ = 0,113

3

3,002

0,5

  • 4

-4,004

Δ = −0,082

-2,997

-3

0,085

Δ = 0,352

0,648

Step 5: Repeat Step 1 to Step 4 for all train vectors

Wael Ouarda - CRNS

36

18

08/11/2021

Δ = 0,352

0,648

Train Database

4. Neural Network: Theory and Application

Train Vector = [2 , -1] ; Train Label = 1

Logistic function =

𝟏

𝟏(cid:2878) 𝒆(cid:3127)𝒙

Step 4: Weights’ Update

Epoch

2

-1

0,506

0,5

Δ = 0,029

𝟏, 𝟎𝟎𝟐

1

Δ = 0,041

0,378

0,876

1,031

1

−𝟎, 𝟗𝟕𝟕

-1

1,5

𝟏, 𝟒𝟗𝟕

  • 2

−𝟐, 𝟎𝟏𝟏

-𝟏, 𝟎𝟎𝟑

-1

Δ = 0,113

3

3,002

0,5

  • 4

-4,004

Δ = −0,082

-2,997

-3

0,085

Step 5: Repeat Step 1 to Step 4 for all train vectors

Wael Ouarda - CRNS

37

5. Naïve Bayes: Theory and Application

Activity

Origin

Stolen

Color

Color

Red

Red

Red

Yellow

Type

Sport

Sport

Sport

Sport

Domicile

Domicile

Domicile

Domicile

Yellow

Sport

Importation

Yellow

Classic

Importation

Yellow Classic

Importation

Yellow

Classic

Advertisement

Domicile

Red

Red

Classic

Importation

Sport

Importation

Yes

No

Yes

No

Yes

No

Yes

No

No

Yes

Learning

𝑷 𝑹𝒆𝒅/𝒀𝒆𝒔 =

𝑷 𝑹𝒆𝒅/𝑵𝒐 =

𝟑

𝟓

𝟐

𝟓

𝑷 𝒀𝒆𝒍𝒍𝒐𝒘/𝒀𝒆𝒔 =

𝑷 𝒀𝒆𝒍𝒍𝒐𝒘/𝑵𝒐 =

𝑷 𝑺𝒑𝒐𝒓𝒕/𝒀𝒆𝒔 =

𝑷 𝑺𝒑𝒐𝒓𝒕/𝑵𝒐 =

𝟒

𝟓

𝟐

𝟓

𝑷 𝑪𝒍𝒂𝒔𝒔𝒊𝒄/𝒀𝒆𝒔 =

𝑷 𝑪𝒍𝒂𝒔𝒔𝒊𝒄/𝑵𝒐 =

𝟐

𝟓

𝟑

𝟓

𝟏

𝟓

𝟑

𝟓

𝑷 𝒀𝒆𝒔 =

𝑷 𝑵𝒐 =

𝟓

𝟏𝟎

𝟓

𝟏𝟎

Type

Origin

𝑷 𝑫𝒐𝒎𝒊𝒄𝒊𝒍𝒆/𝒀𝒆𝒔 =

𝑷 𝑫𝒐𝒎𝒊𝒄𝒊𝒍𝒆/𝑵𝒐 =

𝟐

𝟓

𝟑

𝟓

𝑷 𝑰𝒎𝒑𝒐𝒓𝒕𝒂𝒕𝒊𝒐𝒏/𝒀𝒆𝒔 =

𝑷 𝑰𝒎𝒑𝒐𝒓𝒕𝒂𝒕𝒊𝒐𝒏/𝑵𝒐 =

𝟑

𝟓

𝟐

𝟓

Wael Ouarda - CRNS

38

19

08/11/2021

5. Naïve Bayes: Theory and Application

Sample X= <Red, Classic, Domicile>

𝑷 𝑿/𝒀𝒆𝒔 = 𝑷 𝑹𝒆𝒅/𝒀𝒆𝒔 x 𝑷 𝑪𝒍𝒂𝒔𝒔𝒊𝒄/𝒀𝒆𝒔 x 𝑷 𝑫𝒐𝒎𝒊𝒄𝒊𝒍𝒆/𝒀𝒆𝒔 x 𝑷 𝒀𝒆𝒔

=

𝟑

𝟓

𝟏

𝟓

𝟐

𝟓

𝟓

𝟏𝟎

𝑷 𝑿/𝑵𝒐 = 𝑷 𝑹𝒆𝒅/𝑵𝒐 x 𝑷 𝑪𝒍𝒂𝒔𝒔𝒊𝒄/𝑵𝒐 x 𝑷 𝑫𝒐𝒎𝒊𝒄𝒊𝒍𝒆/𝑵𝒐 x 𝑷 𝑵𝒐

=

𝟐

𝟓

𝟑

𝟓

𝟑

𝟓

𝟓

𝟏𝟎

Testing

𝑷 𝑵𝒐 =

𝟓

𝟏𝟎

𝑷 𝒀𝒆𝒔 =

𝟓

𝟏𝟎

Color

Type

Origin

𝑷 𝑹𝒆𝒅/𝒀𝒆𝒔 =

𝑷 𝑹𝒆𝒅/𝑵𝒐 =

𝟑

𝟓

𝟐

𝟓

𝑷 𝒀𝒆𝒍𝒍𝒐𝒘/𝒀𝒆𝒔 =

𝑷 𝒀𝒆𝒍𝒍𝒐𝒘/𝑵𝒐 =

𝑷 𝑺𝒑𝒐𝒓𝒕/𝒀𝒆𝒔 =

𝑷 𝑺𝒑𝒐𝒓𝒕/𝑵𝒐 =

𝟒

𝟓

𝟐

𝟓

𝑷 𝑪𝒍𝒂𝒔𝒔𝒊𝒄/𝒀𝒆𝒔 =

𝑷 𝑪𝒍𝒂𝒔𝒔𝒊𝒄/𝑵𝒐 =

𝟐

𝟓

𝟑

𝟓

𝟏

𝟓

𝟑

𝟓

𝑷 𝑫𝒐𝒎𝒊𝒄𝒊𝒍𝒆/𝒀𝒆𝒔 =

𝑷 𝑫𝒐𝒎𝒊𝒄𝒊𝒍𝒆/𝑵𝒐 =

𝟐

𝟓

𝟑

𝟓

𝑷 𝑰𝒎𝒑𝒐𝒓𝒕𝒂𝒕𝒊𝒐𝒏/𝒀𝒆𝒔 =

𝑷 𝑰𝒎𝒑𝒐𝒓𝒕𝒂𝒕𝒊𝒐𝒏/𝑵𝒐 =

𝟑

𝟓

𝟐

𝟓

Wael Ouarda - CRNS

39

6. Support Vector Machines: Theory and Application

Basic Idea: Find the appropriate Support Vector which maximize Margin Distance

Class A

Features Space

M1

M2

M1 + M2 = Margin Distance

Class B

Support Vector SV: A.X + B

Wael Ouarda - CRNS

40

20

6. Support Vector Machines: Theory and Application

Basic Idea: Find the appropriate Support Vector which maximize Margin Distance

Case 1: Linear Separation

Case 2: Non Linear Separation

Class A

M1

M2

Class A

Class B

Class B

Support Vector SV: A.X + B

Support Vector SV: F(X)

F is a non linear Function

Wael Ouarda - CRNS

41

6. Support Vector Machines: Theory and Application

Basic Idea: Find the appropriate Support Vector which maximize Margin Distance

Case 1: Linear Separation

Case 2: Non Linear Separation

Class A

M1

M2

Class A

Class B

Class B

Support Vector SV: A.X + B

Support Vector SV: F(X)

F is a non linear Function

Wael Ouarda - CRNS

42

08/11/2021

21

08/11/2021

6. Support Vector Machines: Theory and Application

Features Space

Class A

New Features Space

l

e

n

r

e

K

Class B

Class D

Class C

Wael Ouarda - CRNS

43

6. Support Vector Machines: Theory and Application

Kernel

1

4

2

3

Wael Ouarda - CRNS

44

22

6. Support Vector Machines: Theory and Application

Case 1: Linear Separation

1. K Support Vectors = N-1 where N is the number of samples (K=3)

2.

Initialize 3 linear support vectors

• D1= A1*X+B1

• D2= A2*X+B2

• D3= A3*X+B3

3. Compute the accuracy of SVs

• D1 (75%)

• D2 (100%)

• D3 (100%)

4. Thresholding Acc>85%

• D2 (100%)

• D3 (100%)

4. Compare Margin Distance

• M2 (100%)

• M3 (100%)

• We keep the highest one

Wael Ouarda - CRNS

45

6. Support Vector Machines: Theory and Application

(0,1)

(-1,1)

(0,-1)

(F(0)=1,F(1)=1)

(F(1)=1,F(1)=1)

(F(0)=1,F(-1)=1)

(F(-1)=1,F(1)=1)

Case 2: Non Linear Separation

1. Use a Kernel Function F=x^2

2. Transform Vectors using F function

3. Apply Linear separability

4. K Support Vectors = N-1 where N is the number of samples (K=3)

5.

Initialize 3 linear support vectors

• D1= A1*X+B1

• D2= A2*X+B2

• D3= A3*X+B3

3. Compute the accuracy of SVs

• D1 (75%)

• D2 (100%)

• D3 (100%)

4. Thresholding Acc>85%

• D2 (100%)

• D3 (100%)

4. Compare Margin Distance

• M2 (100%)

• M3 (100%)

• We keep the highest one

Wael Ouarda - CRNS

46

08/11/2021

23

8. How to evaluate the Machine Learning Performance

65

35

𝐹1 − 𝑠𝑐𝑜𝑟𝑒 = 2 ∗

𝑅𝑒𝑐𝑎𝑙𝑙 ∗ 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛

𝑅𝑒𝑐𝑎𝑙𝑙 + 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛

𝑔(cid:3040)(cid:3032)(cid:3028)(cid:3041) =

(𝑅𝑒𝑐𝑎𝑙𝑙 ∗ 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛)

Wael Ouarda - CRNS

47

08/11/2021

24