Deep Learning and Applications
Course 3
Amaury Habrard
Laboratoire Hubert Curien, UMR CNRS 5516
Universit´e Jean Monnet Saint-´Etienne
Semester 1
Today’s content
I Tour on convolutions
I Models for sequence modeling
I Some Problems/Applications
I Recurrent Neural Networks (RNN)
I Long-Short Term Memory networks (LSTM) and Gated
Recurrent Unit cells.
I Bi-directional RNN
I A note on Word embeddings
I Attention and Transformers
Credits: we use the slides/resources of F. Fleuret, K. Gimpel, T.-N.
Le, C. Verloot, S. Wiegre↵e
A tour on convolutions and
transposed convolutions
Vinvent Dumoulin, Francesco Visin. A guide to convolution
arithmetic for deep learning. arXiv:1603.07285, 2020.
https://arxiv.org/abs/1603.07285
https://github.com/vdumoulin/conv_arithmetic
At the end of the section, some images come from
https://www.machinecurve.com/index.php/2019/09/29/
understanding-transposed-convolutions/
Some notations
I n: number of output feature maps
I m: number of input feature maps
I kj kernel size along dimension j
I ij : input size along dimension j
I oj : input size along dimension j
I sj : stride (distance between two consecutive positions of the
kernel) along dimension j
I pj : zero padding (number of zeros concatenated at the
beginning and at the end of an axis) along dimension j
I N: number of dimensions of the kernel
Example of a discrete convolution
0
2
0
3
0
3
2
2
3
0
3
2
2
0
2
0
3
0
3
2
2
0
2
0
1
2
1
3
0
1
0
0
3
0
1
0
0
1
2
1
3
0
1
0
0
1
2
1
2
0
2
2
1
2
0
0
2
1
2
0
0
2
0
2
2
1
2
0
0
2
0
2
1
3
2
2
0
1
3
2
2
0
1
3
2
2
0
0
1
3
2
1
0
1
3
2
1
0
1
3
2
1
12.0
12.0
17.0
10.0
17.0
19.0
9.0
6.0
14.0
12.0
12.0
17.0
10.0
17.0
19.0
9.0
6.0
14.0
12.0
12.0
17.0
10.0
17.0
19.0
9.0
6.0
14.0
3
0
3
2
2
3
0
3
2
2
3
0
3
2
2
0
2
0
3
0
1
0
0
3
0
1
0
0
0
2
0
3
0
1
0
0
0
2
0
1
2
1
2
1
2
0
0
2
1
2
0
0
1
2
1
2
1
2
0
0
1
2
1
2
0
2
1
3
2
2
0
1
3
2
2
0
2
0
2
1
3
2
2
0
2
0
2
0
1
3
2
1
0
1
3
2
1
0
1
3
2
1
12.0
12.0
17.0
10.0
17.0
19.0
9.0
6.0
14.0
12.0
12.0
17.0
10.0
17.0
19.0
9.0
6.0
14.0
12.0
12.0
17.0
10.0
17.0
19.0
9.0
6.0
14.0
3
0
3
2
2
3
0
3
2
2
3
0
3
2
2
3
0
1
0
0
3
0
1
0
0
3
0
1
0
0
0
2
0
2
1
2
0
0
2
1
2
0
0
0
2
0
2
1
2
0
0
0
2
0
1
2
1
1
3
2
2
0
1
3
2
2
0
1
2
1
1
3
2
2
0
1
2
1
2
0
2
0
1
3
2
1
0
1
3
2
1
2
0
2
0
1
3
2
1
2
0
2
12.0
12.0
17.0
10.0
17.0
19.0
9.0
6.0
14.0
12.0
12.0
17.0
10.0
17.0
19.0
9.0
6.0
14.0
12.0
12.0
17.0
10.0
17.0
19.0
9.0
6.0
14.0
3, N = 2,
Figure 1.1: Computing the output values of a discrete convolution.
i2 = 5
o2 = 3
5, m = 9, o1 ⇥
3, s1 = s2 = 1, p1 = p2 = 0
⇤
⇤
n = 25, i1 ⇥
k2 = 3
k1 ⇥
⇥
6.0
17.0
3.0
8.0
17.0
13.0
6.0
4.0
4.0
6.0
17.0
3.0
8.0
17.0
13.0
6.0
4.0
4.0
0
2
0
1
2
1
2
0
2
6.0
17.0
3.0
8.0
17.0
13.0
6.0
4.0
4.0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
2
0
1
2
1
2
0
2
0
3
0
3
2
2
0
0
3
0
3
2
2
0
0
3
0
3
2
2
0
0
3
0
1
0
0
0
0
3
0
1
0
0
0
0
3
0
1
0
0
0
0
2
1
2
0
0
0
0
2
1
2
0
0
0
0
2
1
2
0
0
0
0
2
0
1
2
1
2
0
2
0
1
3
2
2
0
0
0
1
3
2
2
0
0
0
1
3
2
2
0
0
0
0
1
3
2
1
0
0
0
1
3
2
1
0
0
0
1
3
2
1
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
2
0
1
2
1
2
0
Publicité
2
6.0
17.0
3.0
8.0
17.0
13.0
6.0
4.0
4.0
0
2
0
1
2
1
2
0
2
6.0
17.0
3.0
8.0
17.0
13.0
6.0
4.0
4.0
0
2
0
1
2
1
2
0
2
6.0
17.0
3.0
8.0
17.0
13.0
6.0
4.0
4.0
6.0
17.0
3.0
8.0
17.0
13.0
6.0
4.0
4.0
6.0
17.0
3.0
8.0
17.0
13.0
6.0
4.0
4.0
6.0
17.0
3.0
8.0
17.0
13.0
6.0
4.0
4.0
0
2
0
1
2
1
2
0
2
Figure 1.2: Computing the output values of a discrete convolution for N = 2,
i1 = i2 = 5, k1 = k2 = 3, s1 = s2 = 2, and p1 = p2 = 1.
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
3
0
3
2
2
0
0
3
0
3
2
2
0
0
3
0
3
2
2
0
0
3
0
1
0
0
0
0
3
0
1
0
0
0
0
3
0
1
0
0
0
0
2
1
2
0
0
0
0
2
1
2
0
0
0
0
2
1
2
0
0
0
0
1
3
2
2
0
0
0
1
3
2
2
0
0
0
1
3
2
2
0
0
0
0
1
3
2
1
0
0
0
1
3
2
1
0
0
0
1
3
2
1
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
3
0
3
2
2
0
0
3
0
3
2
2
0
0
3
0
3
2
2
0
0
2
0
1
2
1
2
0
2
0
3
0
1
0
0
0
0
3
0
1
0
0
0
0
3
0
1
0
0
0
0
2
1
2
0
0
0
0
2
1
2
0
0
0
0
2
1
2
0
0
0
0
1
3
2
2
0
0
0
1
3
2
2
0
0
0
1
3
2
2
0
0
0
2
0
1
2
1
2
0
2
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
1
3
2
1
0
0
0
1
3
2
1
0
0
0
1
3
2
1
0
7
Figure 1.1: Computing the output values of a discrete convolution.
0
2
0
3
0
3
2
2
3
0
3
2
2
0
2
0
3
0
3
2
2
0
2
0
1
2
1
3
0
1
0
0
3
0
1
0
0
1
2
1
3
0
1
0
0
1
2
1
2
0
2
2
1
2
0
0
2
1
2
0
0
2
0
2
2
1
2
0
0
2
0
2
1
3
2
2
0
1
3
2
2
0
1
3
2
2
0
0
1
3
2
1
0
1
3
2
1
0
1
3
2
1
12.0
12.0
17.0
10.0
17.0
19.0
9.0
6.0
14.0
12.0
12.0
17.0
10.0
17.0
19.0
9.0
6.0
14.0
12.0
12.0
17.0
10.0
17.0
19.0
9.0
6.0
14.0
3
0
3
2
2
3
0
3
2
2
3
0
3
2
2
0
2
0
3
0
1
0
0
3
0
1
0
0
0
2
0
3
0
1
0
0
0
2
0
1
2
1
2
1
2
0
0
2
1
2
0
0
1
2
1
2
1
2
0
0
1
2
1
2
0
2
1
3
2
2
0
1
3
2
2
0
2
0
2
1
3
2
2
0
2
0
2
0
1
3
2
1
0
1
3
2
1
0
1
3
2
1
12.0
12.0
17.0
10.0
17.0
19.0
9.0
6.0
14.0
12.0
12.0
17.0
10.0
17.0
19.0
9.0
6.0
14.0
12.0
Publicité
12.0
17.0
10.0
17.0
19.0
9.0
6.0
14.0
Example of a discrete convolution (2)
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
2
0
0
2
0
0
2
0
0
3
0
3
2
2
0
0
3
0
3
2
2
0
0
3
0
3
2
2
0
1
2
1
1
2
1
1
2
1
0
3
0
1
0
0
0
0
3
0
1
0
0
0
0
3
0
1
0
0
0
2
0
2
2
0
2
2
0
2
0
2
1
2
0
0
0
0
2
1
2
0
0
0
0
2
1
2
0
0
0
0
1
3
2
2
0
0
0
1
3
2
2
0
0
0
1
3
2
2
0
0
0
0
1
3
2
1
0
0
0
1
3
2
1
0
0
0
1
3
2
1
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
6.0
17.0
3.0
8.0
17.0
13.0
6.0
4.0
4.0
6.0
17.0
3.0
8.0
17.0
13.0
6.0
4.0
4.0
6.0
17.0
3.0
8.0
17.0
13.0
6.0
4.0
4.0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
3
0
3
2
2
0
0
3
0
3
2
2
0
0
3
0
3
2
2
0
0
3
0
1
0
0
0
0
3
0
1
0
0
0
0
3
0
1
0
0
0
0
2
0
0
2
0
0
2
0
0
2
1
2
0
0
0
0
2
1
2
0
0
0
0
2
1
2
0
0
0
1
2
1
1
2
1
1
2
1
0
1
3
2
2
0
0
0
1
3
2
2
0
0
0
1
3
2
2
0
0
2
0
2
2
0
2
2
0
2
0
0
1
3
2
1
0
0
0
1
3
2
1
0
0
0
1
3
2
1
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
6.0
17.0
3.0
8.0
17.0
13.0
6.0
4.0
4.0
6.0
17.0
3.0
8.0
17.0
13.0
6.0
4.0
4.0
6.0
17.0
3.0
8.0
17.0
13.0
6.0
4.0
4.0
3
0
3
2
2
3
0
3
2
2
3
0
3
2
2
3
0
1
0
0
3
0
1
0
0
3
0
1
0
0
0
2
0
2
1
2
0
0
2
1
2
0
0
0
2
0
2
1
2
0
0
0
2
0
1
2
1
1
3
2
2
0
1
3
2
2
0
1
2
1
1
3
2
2
0
1
2
1
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
3
0
3
2
2
0
0
3
0
3
2
2
0
0
3
0
3
2
2
0
0
3
0
1
0
0
0
0
3
0
1
0
0
0
0
3
0
1
0
0
0
0
2
1
2
0
0
0
0
2
1
2
0
0
0
0
2
1
2
0
0
0
0
1
3
2
2
0
0
0
1
3
2
2
0
0
0
1
3
2
2
0
0
0
2
0
0
2
0
0
2
0
2
0
2
0
1
3
2
1
0
1
3
2
1
2
0
2
0
1
3
2
1
2
0
2
0
0
1
3
2
1
0
0
0
1
3
2
1
0
0
0
1
3
2
1
0
Publicité
1
2
1
1
2
1
1
2
1
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
2
0
2
2
0
2
2
0
2
12.0
12.0
17.0
10.0
17.0
19.0
9.0
6.0
14.0
12.0
12.0
17.0
10.0
17.0
19.0
9.0
6.0
14.0
12.0
12.0
17.0
10.0
17.0
19.0
9.0
6.0
14.0
6.0
17.0
3.0
8.0
17.0
13.0
6.0
4.0
4.0
6.0
17.0
3.0
8.0
17.0
13.0
6.0
4.0
4.0
6.0
17.0
3.0
8.0
17.0
13.0
6.0
4.0
4.0
3, N = 2,
n = 25, i1 ⇥
k2 = 3
k1 ⇥
⇥
i2 = 5
o2 = 3
5, m = 9, o1 ⇥
3, s1 = s2 = 2, p1 = p2 = 1
⇤
⇤
Figure 1.2: Computing the output values of a discrete convolution for N = 2,
i1 = i2 = 5, k1 = k2 = 3, s1 = s2 = 2, and p1 = p2 = 1.
7
Simplification
I 2-D discrete convolutions (N = 2)
I square inputs i1 = i2 = i
I square kernel size k1 = k2 = k
I same stride along both axes (s1 = s2 = s)
I same zero padding along both axes (p1 = p2 = p)
However, the results can be generalized for the N-D and
non-square cases.
No Zero padding, unit strides
for or any i, k, s = 1 and p = 0: o = (i
On Figure i = 4, k = 3, s = 1, p = 0
Figure 2.1:
input using unit strides (i.e., i = 4, k = 3, s = 1 and p = 0).
(No padding, unit strides) Convolving a 3
k) + 1
⇥
3 kernel over a 4
4
⇥
Figure 2.2:
(Arbitrary padding, unit strides) Convolving a 4
4 kernel over a
5 input padded with a 2
2 border of zeros using unit strides (i.e., i = 5,
⇥
5
⇥
k = 4, s = 1 and p = 2).
⇥
Figure 2.3: (Half padding, unit strides) Convolving a 3
3 kernel over a 5
input using half padding and unit strides (i.e., i = 5, k = 3, s = 1 and p = 1).
5
⇥
⇥
Figure 2.4:
(Full padding, unit strides) Convolving a 3
3 kernel over a 5
input using full padding and unit strides (i.e., i = 5, k = 3, s = 1 and p = 2).
5
⇥
⇥
14
Zero padding, unit strides
Figure 2.1:
input using unit strides (i.e., i = 4, k = 3, s = 1 and p = 0).
(No padding, unit strides) Convolving a 3
⇥
3 kernel over a 4
4
⇥
for any i, k, p, s = 1 o = (i
Figure 2.2:
On figure: i = 5, k = 4, s = 1, p = 2
5
⇥
k = 4, s = 1 and p = 2).
5 input padded with a 2
(Arbitrary padding, unit strides) Convolving a 4
4 kernel over a
2 border of zeros using unit strides (i.e., i = 5,
k) + 2p + 1
⇥
⇥
Figure 2.3: (Half padding, unit strides) Convolving a 3
3 kernel over a 5
input using half padding and unit strides (i.e., i = 5, k = 3, s = 1 and p = 1).
5
⇥
⇥
Figure 2.4:
(Full padding, unit strides) Convolving a 3
3 kernel over a 5
input using full padding and unit strides (i.e., i = 5, k = 3, s = 1 and p = 2).
5
⇥
⇥
14
Figure 2.1:
(No padding, unit strides) Convolving a 3
3 kernel over a 4
input using unit strides (i.e., i = 4, k = 3, s = 1 and p = 0).
⇥
4
⇥
Half (same) padding
Figure 2.2:
5
k = 4, s = 1 and p = 2).
5 input padded with a 2
⇥
⇥
(Arbitrary padding, unit strides) Convolving a 4
4 kernel over a
2 border of zeros using unit strides (i.e., i = 5,
⇥
k/2
for any i, k odd (k = 2l + 1), s = 1, p =
Figure 2.3: (Half padding, unit strides) Convolving a 3
c
⇥
input using half padding and unit strides (i.e., i = 5, k = 3, s = 1 and p = 1).
= l,
⇥
2l = i
3 kernel over a 5
1) = i + 2l
o = i + 2
k/2
c
(k
b
b
5
On figure i = 5, k = 3 and therefore p = 1, s = 1.
This is sometimes referred to as half (or same) padding.
Figure 2.4:
(Full padding, unit strides) Convolving a 3
3 kernel over a 5
input using full padding and unit strides (i.e., i = 5, k = 3, s = 1 and p = 2).
5
⇥
⇥
14
Figure 2.1:
(No padding, unit strides) Convolving a 3
3 kernel over a 4
input using unit strides (i.e., i = 4, k = 3, s = 1 and p = 0).
⇥
4
⇥
Figure 2.2:
(Arbitrary padding, unit strides) Convolving a 4
4 kernel over a
5 input padded with a 2
2 border of zeros using unit strides (i.e., i = 5,
⇥
5
⇥
k = 4, s = 1 and p = 2).
⇥
Half (same) padding
Figure 2.3: (Half padding, unit strides) Convolving a 3
⇥
input using half padding and unit strides (i.e., i = 5, k = 3, s = 1 and p = 1).
3 kernel over a 5
⇥
for any i, k, for p = k
Figure 2.4:
⇥
input using full padding and unit strides (i.e., i = 5, k = 3, s = 1 and p = 2).
(k
(Full padding, unit strides) Convolving a 3
⇥
1) = i + (k
3 kernel over a 5
1 and s = 1
o = i + 2(k
1)
1)
On figure i = 5, k = 3, p = 2, s = 1.
This is sometimes referred to as half (or same) padding.
14
5
5
No zero padding, non unit strides
for any i, k, s for p = 0 and
Figure 2.5: (No zero padding, arbitrary strides) Convolving a 3
a 5
5 input using 2
2 strides (i.e., i = 5, k = 3, s = 2 and p = 0).
+ 1
o =
⇥
⇥
⇥
k
i
3 kernel over
s
b
c
On figure i = 5, k = 3, p = 0, s = 2.
Figure 2.6:
5
⇥
(Arbitrary padding and strides) Convolving a 3
3 kernel over a
2 strides (i.e., i = 5,
⇥
5 input padded with a 1
1 border of zeros using 2
k = 3, s = 2 and p = 1).
⇥
⇥
Figure 2.7:
(Arbitrary padding and strides) Convolving a 3
3 kernel over a
6
⇥
6 input padded with a 1
1 border of zeros using 2
2 strides (i.e., i = 6,
k = 3, s = 2 and p = 1). In this case, the bottom row and right column of the
⇥
⇥
⇥
zero padded input are not covered by the kernel.
(a) The kernel has to slide two steps
(b) The kernel has to slide one step of
to the right to touch the right side of
size two to the right to touch the right
the input (and equivalently downwards).
side of the input (and equivalently down-
Adding one to account for the initial ker-
wards). Adding one to account for the
nel position, the output size is 3
3.
initial kernel position, the output size is
2.
2
Figure 2.8: Counting kernel positions.
17
Zero padding, non unit strides
Figure 2.5: (No zero padding, arbitrary strides) Convolving a 3
a 5
2 strides (i.e., i = 5, k = 3, s = 2 and p = 0).
5 input using 2
⇥
3 kernel over
⇥
⇥
The most general case: for any i, k, p, s
5 input padded with a 1
Figure 2.6:
5
k = 3, s = 2 and p = 1).
(Arbitrary padding and strides) Convolving a 3
i + 2p
s
1 border of zeros using 2
k
⇥
o =
+ 1
⇥
⇥
c
b
3 kernel over a
2 strides (i.e., i = 5,
⇥
On figure i = 5, k = 3, p = 1, s = 2.
Figure 2.7:
(Arbitrary padding and strides) Convolving a 3
3 kernel over a
6
⇥
6 input padded with a 1
1 border of zeros using 2
2 strides (i.e., i = 6,
k = 3, s = 2 and p = 1). In this case, the bottom row and right column of the
⇥
⇥
⇥
zero padded input are not covered by the kernel.
(a) The kernel has to slide two steps
(b) The kernel has to slide one step of
to the right to touch the right side of
size two to the right to touch the right
the input (and equivalently downwards).
side of the input (and equivalently down-
Adding one to account for the initial ker-
wards). Adding one to account for the
nel position, the output size is 3
3.
initial kernel position, the output size is
2.
2
Figure 2.8: Counting kernel positions.
17
Pooling arithmetic
The most general case :for any i, k, and, s
Figure 2.5: (No zero padding, arbitrary strides) Convolving a 3
a 5
5 input using 2
2 strides (i.e., i = 5, k = 3, s = 2 and p = 0).
+ 1
o =
⇥
⇥
⇥
k
i
3 kernel over
s
b
c
It holds for any type of pooling (note that pooling does not involve
zero padding).
On figure i = 5, k = 3, p = 0, s = 2.
Figure 2.6:
5
k = 3, s = 2 and p = 1).
5 input padded with a 1
⇥
(Arbitrary padding and strides) Convolving a 3
1 border of zeros using 2
⇥
⇥
3 kernel over a
2 strides (i.e., i = 5,
⇥
Figure 2.7:
(Arbitrary padding and strides) Convolving a 3
3 kernel over a
6
⇥
6 input padded with a 1
1 border of zeros using 2
2 strides (i.e., i = 6,
k = 3, s = 2 and p = 1). In this case, the bottom row and right column of the
⇥
⇥
⇥
zero padded input are not covered by the kernel.
(a) The kernel has to slide two steps
(b) The kernel has to slide one step of
to the right to touch the right side of
size two to the right to touch the right
the input (and equivalently downwards).
side of the input (and equivalently down-
Adding one to account for the initial ker-
wards). Adding one to account for the
nel position, the output size is 3
3.
initial kernel position, the output size is
2.
2
Figure 2.8: Counting kernel positions.
17
Transposed convolutions
We use the slides of F. Fleuret
1
2
1
0
1
2
-1
0
1
2
-1
0
2
1
-1
2
1
0
-1
0
1
2
-1
0
2
-1
0
-1
w
w
w
w
w
w
w
9
0
1
3
-5
-3
6
Output
W
w + 1
Convolution layer
Input
1
4
-1
0
2
-2
1
3
3
1
W
Kernel
1
2
0
-1
w
Fran¸cois Fleuret
EE-559 – Deep learning / 7.1. Transposed convolutions
5 / 14
1
2
1
-1
0
2
-1
0
-1
2
1
0
1
Kernel
-1
0
1
2
0
-1
0
1
2
2
-1
1
0
2
-1
w
w
w
w
w
w
w
0
1
3
-5
-3
6
1
1
Convolution layer
Input
4
-1
0
2
-2
1
3
3
1
W
2
0
-1
w
9
Output
W
w + 1
Fran¸cois Fleuret
EE-559 – Deep learning / 7.1. Transposed convolutions
5 / 14
1
2
0
1
-1
0
2
-1
0
-1
-1
1
2
1
Kernel
-1
1
0
2
0
2
1
0
2
-1
1
0
2
-1
w
w
w
w
w
w
w
1
3
-5
-3
6
Convolution layer
Input
1
Publicité
4
-1
0
2
-2
1
3
3
1
W
1
2
0
-1
w
9
0
Output
W
w + 1
Fran¸cois Fleuret
EE-559 – Deep learning / 7.1. Transposed convolutions
5 / 14
1
2
1
0
2
-1
0
2
-1
0
-1
Kernel
-1
1
0
1
-1
2
1
2
2
1
0
0
-1
2
1
0
-1
w
w
w
w
w
w
w
3
-5
-3
6
Convolution layer
Input
1
4
-1
0
2
-2
1
3
3
1
W
1
2
0
-1
w
Output
9
0
1
W
w + 1
Fran¸cois Fleuret
EE-559 – Deep learning / 7.1. Transposed convolutions
5 / 14
1
2
1
2
1
0
-1
2
0
-1
0
-1
-1
0
2
1
Kernel
-1
2
1
0
-1
1
0
2
2
1
0
-1
w
w
w
w
w
w
w
-5
-3
6
Convolution layer
Input
1
4
-1
0
2
-2
1
3
3
1
W
1
2
0
-1
w
Output
9
0
1
3
W
w + 1
Fran¸cois Fleuret
EE-559 – Deep learning / 7.1. Transposed convolutions
5 / 14
1
2
1
0
1
2
0
2
-1
0
-1
-1
0
2
1
1
Kernel
-1
1
0
0
-1
2
0
2
-1
1
2
-1
w
w
w
w
w
w
w
-3
6
Convolution layer
Input
1
4
-1
0
2
-2
1
3
3
1
W
1
2
0
-1
w
Output
9
0
1
3
-5
W
w + 1
Fran¸cois Fleuret
EE-559 – Deep learning / 7.1. Transposed convolutions
5 / 14
1
2
1
2
1
0
-1
2
0
-1
-1
2
1
0
1
Kernel
-1
0
2
0
-1
2
1
0
2
-1
0
1
-1
w
w
w
w
w
w
w
6
Convolution layer
Input
1
4
-1
0
2
-2
W
1
Output
1
2
3
3
1
0
-1
w
9
0
1
3
-5
-3
W
w + 1
Fran¸cois Fleuret
EE-559 – Deep learning / 7.1. Transposed convolutions
5 / 14
1
2
1
2
1
0
-1
2
1
0
1
Kernel
-1
0
1
2
0
-1
2
1
0
2
-1
0
2
-1
-1
0
-1
w
w
w
w
w
w
w
Convolution layer
Input
1
4
-1
0
2
-2
W
Output
1
1
3
2
3
1
0
-1
w
9
0
1
3
-5
-3
6
W
w + 1
Fran¸cois Fleuret
EE-559 – Deep learning / 7.1. Transposed convolutions
5 / 14
1
2
1
2
1
0
-1
0
2
1
-1
0
1
2
-1
0
1
2
-1
0
1
2
-1
0
2
-1
0
-1
w
w
w
w
w
w
w
Convolution layer
Input
1
4
-1
0
2
-2
1
3
3
1
W
Kernel
1
2
0
-1
w
Output
9
0
1
3
-5
-3
6
W
w + 1
Fran¸cois Fleuret
EE-559 – Deep learning / 7.1. Transposed convolutions
5 / 14
+
1
2
2
1
4
3
-1
2
1
-2
6
0
-1
2
1
-1
2
-1
-3
0
0
-1
-2
1
2
7
4
-4
-2
1
Output
W + w
1
Transposed convolution layer
Input
2
3
0
-1
W
1
Kernel
2
w
-1
Fran¸cois Fleuret
EE-559 – Deep learning / 7.1. Transposed convolutions
6 / 14
1
2
1
-1
2
1
-1
2
-1
1
3
-1
-3
Kernel
2
w
6
0
0
0
-1
-2
1
7
4
-4
-2
1
Transposed convolution layer
Input
3
0
-1
W
-1
-2
2
2
4
Output
W + w
1
1
2
2
+
Fran¸cois Fleuret
EE-559 – Deep learning / 7.1. Transposed convolutions
6 / 14
1
2
-1
1
2
1
-1
2
-1
1
-1
Kernel
2
w
0
0
0
-1
-2
1
4
-4
-2
1
Transposed convolution layer
Input
3
0
-1
W
2
-1
-2
6
-3
2
1
4
3
2
+
2
7
Output
W + w
1
Fran¸cois Fleuret
EE-559 – Deep learning / 7.1. Transposed convolutions
6 / 14
1
2
1
-1
2
-1
1
2
-1
1
-1
Kernel
2
w
-1
-2
1
-4
-2
1
Transposed convolution layer
2
4
3
2
Input
W
0
-1
2
-1
-3
0
0
3
1
-2
6
0
Output
2
7
4
W + w
1
+
Fran¸cois Fleuret
EE-559 – Deep learning / 7.1. Transposed convolutions
6 / 14
1
2
1
-1
2
1
-1
2
-1
1
-1
Kernel
<...