Release of Tesseract 5.1 Text Recognition System
The release of the optical character recognition system Tesseract 5.1 has been published, supporting UTF-8 character recognition and texts in over 100 languages, including Russian, Kazakh, Belarusian, and Ukrainian. The results can be saved as plain text or in HTML (hOCR), ALTO (XML), PDF, and TSV formats. The system was originally developed between 1985 and 1995 in a laboratory at Hewlett Packard, in […]
