The release of the Tesseract 5.4.0 text recognition system

The release of the optical character recognition system Tesseract 5.4.0 has been announced, supporting UTF-8 character recognition and texts in over 100 languages, including Russian, Kazakh, Belarusian, and Ukrainian. Results can be saved as plain text as well as in HTML (hOCR), ALTO (XML), PDF, and TSV formats. Initially, the system was developed between 1985 and 1995 at the Hewlett Packard lab, and in 2005 the code was open-sourced under the Apache license, subsequently evolving with the participation of Google employees. The project's source texts are distributed under the Apache 2.0 license.

Tesseract includes a console utility and the libtesseract library for embedding text recognition features into other applications. Notable third-party GUI interfaces supporting Tesseract include gImageReader, VietOCR, and YAGF. It offers two recognition engines: the classic one, which recognizes text by matching individual character templates, and a new one based on a recurrent neural network (LSTM) machine learning system, optimized for recognizing entire lines of text and significantly improving accuracy. Pre-trained models are available for 123 languages. Performance optimization modules using OpenMP and SIMD instructions such as AVX2, AVX, AVX512F, NEON, or SSE4.1 are provided.

Key Improvements:

  • Support for rendering and exporting in PAGE-XML format has been added.
  • The ability to train models using PNG files instead of LSTMF files has been implemented.
  • PDF rendering has been improved.
  • The API for text tilt detection has been expanded.
  • Performance issues identified during scanning in the Coverity system have been resolved.

Source: opennet.ru

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster