7. Building a Multilingual OCR Engine
Training LSTM networks on 100 languages and test results
Ray Smith, Google Inc.
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
Internationalization: The Convex Hull of Languages
If you want to develop multilingual OCR, work on these first...
29k possible graphemes
of 7 or more unicodes
English
The most worked-on language - the hardest to beat
Kannada
Vietnamese
Lot of unusual diacritics
Multiple unicodes
combine into ligatures
Hindi
Russian
Case ambiguities
Thai
Japanese
Arabic
Right-to-left, Joined
characters and Bidi (bi-
directional)
Stacking diacritics,
ambiguous characters
4 different scripts written
horizontally and vertically
on the same page
Urdu
Until very recently, not even
machine renderable!
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
Bidirectional issues
3
2
1
U+0028 - open parenthesis
U+0029 - close parenthesis
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
What’s a “character” in Devanagari?
Result
Unicode
0930
0930 094d
0930 094d 0926
0930 094d 0926 094d
0930 094d 0926 094d 0935
0930 094d 0926 094d 0935 093f
0930 094d 0926 094d 0935 093f 0915
Transliteration
ra
r
rda
rd
rdva
rdvi
rdvika
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
What’s a “character” in Kannada?
Result
Unicode
Transliteration
ರ
ದ
ವ
(cid:349)
ದ(cid:143)
(cid:340)(cid:143)
ದ(cid:353)(cid:143)
ದ(cid:353)್(cid:143)
ದ(cid:353)(cid:180)(cid:143)
(cid:296)(cid:353)(cid:180)(cid:143)
0cb0
0ca6
0cb5
0cb0 0ccd
0cb0 0ccd 0ca6
0cb0 0ccd 0ca6 0ccd
0cb0 0ccd 0ca6 0ccd 0cb5
0cb0 0ccd 0ca6 0ccd 0cb5 0ccd
0cb0 0ccd 0ca6 0ccd 0cb5 0ccd 0c95
0cb0 0ccd 0ca6 0ccd 0cb5 0ccd 0c95 0cbf
ra
da
va
r
rda
rd
rdva
rdv
rdvka
rdvki
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
Universal Character/Grapheme Encoding/Compression
Extension to Tesseract’s UNICHARSET to make the output Softmax smaller
Alphabetic
NFKC Normalize
Unicharset
Codes
Indic
Hangul
Han
Split into Unicodes
Split into Jamos
Radical-stroke-index codes
Ligatures
Handled
Publicité
Several
Codes
Triple
Codes
Triple
Codes
Compressed
Codes
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
Training International OCR Engines
Tesseract Training
Web Crawl
Repository
Language ID
Map-Reduce
Cleaned Language
Corpora
OCR Engine
Training
Manually generated Files
Eng
OCR Shape Files
Eng
Eng
Realistic Text
Rendering
32 fonts
Language Model Files
Eng
Text Filtration
Dirty Language
Corpora
Language Model
Generation
Eng
Photo By Steve Jurvetson (http://www.flickr.com/photos/jurvetson/162116759) [CC BY 2.0 (http://creativecommons.
org/licenses/by/2.0)], via Wikimedia Commons
https://commons.wikimedia.org/wiki/File%3AHawk_eye.jpg
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
Training International OCR Engines
T-LSTM Training
Web Crawl
Repository
Language ID
Map-Reduce
Cleaned Language
Corpora
OCR Engine
Training
x500
Eng
Realistic Text
Rendering
x500
Manually generated Files
Eng
OCR Shape Files
Eng
x100
5000 fonts
Language Model Files
Eng
Text Filtration
Dirty Language
Corpora
Language Model
Generation
Eng
Photo By Steve Jurvetson (http://www.flickr.com/photos/jurvetson/162116759) [CC BY 2.0 (http://creativecommons.
org/licenses/by/2.0)], via Wikimedia Commons
https://commons.wikimedia.org/wiki/File%3AHawk_eye.jpg
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
Training
● Synthetic Training data:
With bounding boxes
Instead of CTC
● About 500k lines per language
● Random book-like degradation
● (Almost) Same network specification for each language:
● Convergence in 3-5 days or more [G2,0C2,2FT16P3,3LQ1,64L1,128RtL1,128LS1,256]
[G2,0C2,2FT16P3,3LQ1,64L1,128RtL1,128LS1,512]
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
Testing
Testset from Google Books:
● Single Lines cut from older books.
● Hand typed. Accuracy far from perfect.
● 1000 lines * ~50 languages.
Caveat: Does not allow Tesseract to adapt to a whole page. (T-LSTM doesn’t adapt)
Example:
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
Overall effect on 51 Languages
Tesseract 3.04 Baseline
Impossible to resolve
individual language
results, but overall
feel is improved
Notice this annoying
precision ceiling is
completely gone in the
new version
T-LSTM (no dict)
T-LSTM + Dict
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
Accuracy Results: Selection of Latin Languages
Truth
Words Char Error Rates
%Change Word Error Rates
%Change
Base
Publicité
T-LSTM LSTM+Dict
Base T-LSTM LSTM+Dict
Lang
Czech
English
French
19319
21543
20746
Hungarian
17977
Indonesian
16616
Dutch
18878
Norwegian
18129
Portuguese 17726
Spanish
20333
3.08
2.51
5.98
3.75
2.77
6.23
2.19
2.91
4.28
1.86
1.95
3.1
2.95
1.96
6.65
1.74
1.9
2.5
1.81
1.76
2.98
2.86
2.26
5.97
1.81
1.87
2.31
-41.23 12.69
-29.88
7.58
-50.17 16.63
-23.73 14.61
-18.41
9.45
8.26
7.18
11.82
12.26
8.64
7.58
5.77
10.47
11.45
-40.27
-23.88
-37.04
-21.63
7.51
-20.53
-4.17 16.38
17.55
15.75
-3.85
-17.35
7.73
-35.74 10.09
-46.03 11.63
5.87
7.46
8.55
5.7
-26.26
6.75
7.45
-33.10
-35.94
Vietnamese 22472
4.04
3.55
2.61
-35.40 10.89
10.05
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
7.93
-27.18
Accuracy Results: Latin Langs ROC curves
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
Accuracy Results: Cyrillic Languages
Lang
Truth
Words Char Error Rates
%Change Word Error Rates
%Change
Base T-LSTM LSTM+Dict
Base T-LSTM LSTM+Dict
Belarusian
11697
7.72
4.25
Publicité
Bulgarian
18457 19.68
3.78
Macedonian 17117 11.22
3.18
Russian
14993 19.06
4
Ukrainian
17123
6.63
2.77
4.60
3.85
3.08
4.10
3.27
-40.41 20.63
16.66
16.38
-20.60
-80.44 29.88
13.18
12.4
-58.50
-72.55 18.92
11.69
10.61
-43.92
-78.49 30.83
15.34
14.04
-54.46
-50.68 16.78
11.6
11.34
-32.42
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
Accuracy Results (Cyrillic Langs)
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
Accuracy Results (Indic Languages)
Lang
Truth Words Char Error Rates
%Change Word Error Rates
%Change
Base T-LSTM LSTM+Dict
Base
T-
LSTM LSTM+Dict
Bengali
Gujarati
Hindi
16865 19.59
22.33
19.25
-1.74 42.79
42.23
39.13
-8.55
19874
45.7
22.38
18.35
-59.85 56.46
49.46
43.62
-22.74
27539 14.05
13.56
11.68
-16.87 30.92
30.19
26.92
-12.94
Kannada
13673 31.55
13.57
12.12
-61.58
63.2
47.11
40.71
-35.59
Marathi
21486 29.04
12.07
9.18
-68.39 47.19
32.79
25.92
-45.07
Nepalese
20606 29.56
18.38
15.63
-47.12 50.19
42.39
36.59
-27.10
Panjabi
Tamil
Telugu
27651 46.23
19.89
16.1
-65.17 54.89
40.18
37.15
-32.32
Publicité
10033
26.1
10.09
9.3
-64.37 63.58
30.65
26.78
-57.88
14133 31.47
31.51
27.8
-11.66 63.63
63.79
59.36
-6.71
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
Accuracy Results (Indic Langs)
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
Accuracy Results (Other Languages)
Lang
Truth
Words Char Error Rates
%Change Word Error Rates
%Change
Base T-LSTM LSTM+Dict
Base T-LSTM LSTM+Dict
Hebrew
21919 12.29
8.08
Thai
32173 60.52
12.32
8.36
7.73
-31.98 34.28
28.81
29.91
-12.75
-87.23 96.91
38.46
22.70
-76.58
Yiddish
21674 20.76
10.04
10.62
-48.84 56.77
34.32
36.53
-35.65
Farsi
23691 74.87
16.53
16.84
-77.51 65.44
40.43
42.80
-34.60
Chinese(S)
31045
8.25
7.25
6.42
-22.18 10.82
11.41
9.32
-13.86
Chinese(T)
28700 13.22
12.83
10.35
-21.71 18.56
20.28
15.78
-14.98
Japanese
29574 18.65
16.97
11.53
-38.18 31.66
35.49
19.94
-37.02
Korean
25687 31.19
9.62
6.67
-78.61 40.09
18.69
9.82
-75.51
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
Accuracy Results (Other Langs)
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
Conclusion
● Tesseract is surprisingly general
● T-LSTM is much better almost across the board
○
Some Language model integration issues remain
● Some languages remain difficult, but
● Neural networks are taking over!
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece
Thanks for Listening!
Questions?
Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece