Building a Multilingual OCR Engine: Training LSTM Networks on 100 Languages and Test Results

Page 1 sur 21Lecteur de document UniversityLib

Building a Multilingual OCR Engine: Training LSTM Networks on 100 Languages and Test Results

Optical Character Recognition and Machine Learning · notes

Voir tous les documents en intelligence artificielle et données

7. Building a Multilingual OCR Engine

Training LSTM networks on 100 languages and test results

Ray Smith, Google Inc.

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

Internationalization: The Convex Hull of Languages

If you want to develop multilingual OCR, work on these first...

29k possible graphemes

of 7 or more unicodes

English

The most worked-on language - the hardest to beat

Kannada

Vietnamese

Lot of unusual diacritics

Multiple unicodes

combine into ligatures

Hindi

Russian

Case ambiguities

Thai

Japanese

Arabic

Right-to-left, Joined

characters and Bidi (bi-

directional)

Stacking diacritics,

ambiguous characters

4 different scripts written

horizontally and vertically

on the same page

Urdu

Until very recently, not even

machine renderable!

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

Bidirectional issues

3

2

1

U+0028 - open parenthesis

U+0029 - close parenthesis

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

What’s a “character” in Devanagari?

Result

Unicode

0930

0930 094d

0930 094d 0926

0930 094d 0926 094d

0930 094d 0926 094d 0935

0930 094d 0926 094d 0935 093f

0930 094d 0926 094d 0935 093f 0915

Transliteration

ra

r

rda

rd

rdva

rdvi

rdvika

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

What’s a “character” in Kannada?

Result

Unicode

Transliteration

(cid:349)

ದ(cid:143)

(cid:340)(cid:143)

ದ(cid:353)(cid:143)

ದ(cid:353)್(cid:143)

ದ(cid:353)(cid:180)(cid:143)

(cid:296)(cid:353)(cid:180)(cid:143)

0cb0

0ca6

0cb5

0cb0 0ccd

0cb0 0ccd 0ca6

0cb0 0ccd 0ca6 0ccd

0cb0 0ccd 0ca6 0ccd 0cb5

0cb0 0ccd 0ca6 0ccd 0cb5 0ccd

0cb0 0ccd 0ca6 0ccd 0cb5 0ccd 0c95

0cb0 0ccd 0ca6 0ccd 0cb5 0ccd 0c95 0cbf

ra

da

va

r

rda

rd

rdva

rdv

rdvka

rdvki

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

Universal Character/Grapheme Encoding/Compression

Extension to Tesseract’s UNICHARSET to make the output Softmax smaller

Alphabetic

NFKC Normalize

Unicharset

Codes

Indic

Hangul

Han

Split into Unicodes

Split into Jamos

Radical-stroke-index codes

Ligatures

Handled

Publicité

Several

Codes

Triple

Codes

Triple

Codes

Compressed

Codes

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

Training International OCR Engines

Tesseract Training

Web Crawl

Repository

Language ID

Map-Reduce

Cleaned Language

Corpora

OCR Engine

Training

Manually generated Files

Eng

OCR Shape Files

Eng

Eng

Realistic Text

Rendering

32 fonts

Language Model Files

Eng

Text Filtration

Dirty Language

Corpora

Language Model

Generation

Eng

Photo By Steve Jurvetson (http://www.flickr.com/photos/jurvetson/162116759) [CC BY 2.0 (http://creativecommons.

org/licenses/by/2.0)], via Wikimedia Commons

https://commons.wikimedia.org/wiki/File%3AHawk_eye.jpg

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

Training International OCR Engines

T-LSTM Training

Web Crawl

Repository

Language ID

Map-Reduce

Cleaned Language

Corpora

OCR Engine

Training

x500

Eng

Realistic Text

Rendering

x500

Manually generated Files

Eng

OCR Shape Files

Eng

x100

5000 fonts

Language Model Files

Eng

Text Filtration

Dirty Language

Corpora

Language Model

Generation

Eng

Photo By Steve Jurvetson (http://www.flickr.com/photos/jurvetson/162116759) [CC BY 2.0 (http://creativecommons.

org/licenses/by/2.0)], via Wikimedia Commons

https://commons.wikimedia.org/wiki/File%3AHawk_eye.jpg

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

Training

● Synthetic Training data:

With bounding boxes

Instead of CTC

● About 500k lines per language

● Random book-like degradation

● (Almost) Same network specification for each language:

● Convergence in 3-5 days or more [G2,0C2,2FT16P3,3LQ1,64L1,128RtL1,128LS1,256]

[G2,0C2,2FT16P3,3LQ1,64L1,128RtL1,128LS1,512]

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

Testing

Testset from Google Books:

● Single Lines cut from older books.

● Hand typed. Accuracy far from perfect.

● 1000 lines * ~50 languages.

Caveat: Does not allow Tesseract to adapt to a whole page. (T-LSTM doesn’t adapt)

Example:

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

Overall effect on 51 Languages

Tesseract 3.04 Baseline

Impossible to resolve

individual language

results, but overall

feel is improved

Notice this annoying

precision ceiling is

completely gone in the

new version

T-LSTM (no dict)

T-LSTM + Dict

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

Accuracy Results: Selection of Latin Languages

Truth

Words Char Error Rates

%Change Word Error Rates

%Change

Base

Publicité

T-LSTM LSTM+Dict

Base T-LSTM LSTM+Dict

Lang

Czech

English

French

19319

21543

20746

Hungarian

17977

Indonesian

16616

Dutch

18878

Norwegian

18129

Portuguese 17726

Spanish

20333

3.08

2.51

5.98

3.75

2.77

6.23

2.19

2.91

4.28

1.86

1.95

3.1

2.95

1.96

6.65

1.74

1.9

2.5

1.81

1.76

2.98

2.86

2.26

5.97

1.81

1.87

2.31

-41.23 12.69

-29.88

7.58

-50.17 16.63

-23.73 14.61

-18.41

9.45

8.26

7.18

11.82

12.26

8.64

7.58

5.77

10.47

11.45

-40.27

-23.88

-37.04

-21.63

7.51

-20.53

-4.17 16.38

17.55

15.75

-3.85

-17.35

7.73

-35.74 10.09

-46.03 11.63

5.87

7.46

8.55

5.7

-26.26

6.75

7.45

-33.10

-35.94

Vietnamese 22472

4.04

3.55

2.61

-35.40 10.89

10.05

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

7.93

-27.18

Accuracy Results: Latin Langs ROC curves

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

Accuracy Results: Cyrillic Languages

Lang

Truth

Words Char Error Rates

%Change Word Error Rates

%Change

Base T-LSTM LSTM+Dict

Base T-LSTM LSTM+Dict

Belarusian

11697

7.72

4.25

Publicité

Bulgarian

18457 19.68

3.78

Macedonian 17117 11.22

3.18

Russian

14993 19.06

4

Ukrainian

17123

6.63

2.77

4.60

3.85

3.08

4.10

3.27

-40.41 20.63

16.66

16.38

-20.60

-80.44 29.88

13.18

12.4

-58.50

-72.55 18.92

11.69

10.61

-43.92

-78.49 30.83

15.34

14.04

-54.46

-50.68 16.78

11.6

11.34

-32.42

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

Accuracy Results (Cyrillic Langs)

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

Accuracy Results (Indic Languages)

Lang

Truth Words Char Error Rates

%Change Word Error Rates

%Change

Base T-LSTM LSTM+Dict

Base

T-

LSTM LSTM+Dict

Bengali

Gujarati

Hindi

16865 19.59

22.33

19.25

-1.74 42.79

42.23

39.13

-8.55

19874

45.7

22.38

18.35

-59.85 56.46

49.46

43.62

-22.74

27539 14.05

13.56

11.68

-16.87 30.92

30.19

26.92

-12.94

Kannada

13673 31.55

13.57

12.12

-61.58

63.2

47.11

40.71

-35.59

Marathi

21486 29.04

12.07

9.18

-68.39 47.19

32.79

25.92

-45.07

Nepalese

20606 29.56

18.38

15.63

-47.12 50.19

42.39

36.59

-27.10

Panjabi

Tamil

Telugu

27651 46.23

19.89

16.1

-65.17 54.89

40.18

37.15

-32.32

Publicité

10033

26.1

10.09

9.3

-64.37 63.58

30.65

26.78

-57.88

14133 31.47

31.51

27.8

-11.66 63.63

63.79

59.36

-6.71

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

Accuracy Results (Indic Langs)

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

Accuracy Results (Other Languages)

Lang

Truth

Words Char Error Rates

%Change Word Error Rates

%Change

Base T-LSTM LSTM+Dict

Base T-LSTM LSTM+Dict

Hebrew

21919 12.29

8.08

Thai

32173 60.52

12.32

8.36

7.73

-31.98 34.28

28.81

29.91

-12.75

-87.23 96.91

38.46

22.70

-76.58

Yiddish

21674 20.76

10.04

10.62

-48.84 56.77

34.32

36.53

-35.65

Farsi

23691 74.87

16.53

16.84

-77.51 65.44

40.43

42.80

-34.60

Chinese(S)

31045

8.25

7.25

6.42

-22.18 10.82

11.41

9.32

-13.86

Chinese(T)

28700 13.22

12.83

10.35

-21.71 18.56

20.28

15.78

-14.98

Japanese

29574 18.65

16.97

11.53

-38.18 31.66

35.49

19.94

-37.02

Korean

25687 31.19

9.62

6.67

-78.61 40.09

18.69

9.82

-75.51

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

Accuracy Results (Other Langs)

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

Conclusion

● Tesseract is surprisingly general

● T-LSTM is much better almost across the board

Some Language model integration issues remain

● Some languages remain difficult, but

● Neural networks are taking over!

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece

Thanks for Listening!

Questions?

Tesseract Blends Old and New OCR Technology - DAS2016 Tutorial - Santorini - Greece