Logistic Regression
Logistic regression is a learning algorithm used in a supervised learning problem when the output 𝑦 are
all either zero or one. The goal of logistic regression is to minimize the error between its predictions and
training data.
Example: Cat vs No - cat
Given an image represented by a feature vector 𝑥, the algorithm will evaluate the probability of a cat
Publicité
being in that image.
𝐺𝑖𝑣𝑒𝑛 𝑥 , 𝑦̂ = 𝑃(𝑦 = 1|𝑥), where 0 ≤ 𝑦̂ ≤ 1
The parameters used in Logistic regression are:
• The input features vector: 𝑥 ∈ ℝ𝑛𝑥, where 𝑛𝑥 is the number of features
• The training label: 𝑦 ∈ 0,1
• The weights: 𝑤 ∈ ℝ𝑛𝑥, where 𝑛𝑥 is the number of features
Publicité
• The threshold: 𝑏 ∈ ℝ
• The output: 𝑦̂ = 𝜎(𝑤𝑇𝑥 + 𝑏)
• Sigmoid function: s = 𝜎(𝑤𝑇𝑥 + 𝑏) = 𝜎(𝑧)=
1
1+ 𝑒−𝑧
(𝑤𝑇𝑥 + 𝑏) is a linear function (𝑎𝑥 + 𝑏), but since we are looking for a probability constraint between
Publicité
[0,1], the sigmoid function is used. The function is bounded between [0,1] as shown in the graph above.
Some observations from the graph:
•
•
•
If 𝑧 is a large positive number, then 𝜎(𝑧) = 1
Publicité
If 𝑧 is small or large negative number, then 𝜎(𝑧) = 0
If 𝑧 = 0, then 𝜎(𝑧) = 0.5