Release of Tesseract 5.1 Text Recognition System

The release of the Optical Character Recognition system Tesseract 5.1 has been announced, supporting UTF-8 character recognition and texts in over 100 languages, including Russian, Kazakh, Belarusian, and Ukrainian. Results can be saved as plain text or in HTML (hOCR), ALTO (XML), PDF, and TSV formats. Originally developed between 1985 and 1995 in a laboratory at Hewlett Packard, the code was opened under the Apache license in 2005 and has since been further developed with the involvement of Google employees. The source texts of the project are distributed under the Apache 2.0 license.

Tesseract includes a console utility and the libtesseract library for embedding text recognition features in other applications. Among the third-party GUI interfaces supporting Tesseract are gImageReader, VietOCR, and YAGF. It offers two recognition engines: a classic one that recognizes text at the character template level, and a new one based on a recurrent neural network (RNN) machine learning system optimized for recognizing entire lines, allowing for a significant increase in accuracy. Pre-trained models are published for 123 languages. Performance optimization modules use OpenMP and SIMD instructions AVX2, AVX, NEON, or SSE4.1.

Main improvements in Tesseract 5.1:

  • The capability to process areas with images and lines when outputting in ALTO, hOCR, and text formats has been implemented.
  • A new parameter curl_timeout has been added to curl_easy_setopt.
  • The build system has been improved.
  • Efforts have been made to remove unused code.
  • Crashes caused by improper handling of null pointers in the PageIterator::Orientation class have been fixed.

Source: opennet.ru

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster