Keywords: metric books, OCR, paddleOCR, gigaChat-MAX, LLM, pre-reform orthography, CER, WER, TF‑IDF
UDC 681.3
DOI: 10.26102/2310-6018/2026.59.8.002
The paper addresses the problem of improving text recognition quality for metric books of the Russian Empire (late 19th – early 20th centuries), which are a valuable historical source for genealogy and demography. Conventional optical character recognition (OCR) systems trained on modern printed texts achieve an accuracy of no more than 30 % on such documents due to a combination of factors: paper deterioration, ink fading, pre-reform orthography (letters ѣ, і, ѳ, ѵ) and mixing of Cyrillic and Latin glyphs. The study presents a comparative analysis of two approaches: fine-tuning the PaddleOCR model on historical data versus using a pipeline of baseline PaddleOCR with the large language model GigaChat-MAX without fine-tuning. For post-processing, a specialised prompt was developed, including character substitution rules, a dictionary of frequent terms, and rules for restoring pre-reform spelling. Experiments on a sample of 50 spreads from a metric book showed that fine-tuning increases accuracy to 48 %, while the integration with the language model achieves 65 %, with the character error rate reduced from 70 % to 35 %. The language model corrects more than half of all errors, being particularly effective in replacing Latin letters with Cyrillic and restoring historical endings. The scientific novelty lies in the first direct comparison of the two strategies and the demonstration of the advantage of post-processing with a language model without the need for labelled data and computational costs. The practical significance is that the proposed approach enables automated recognition of metric books with 65 % accuracy, reducing manual transcription effort by more than half.
1. Levenshtein V.I. Binary codes capable of correcting deletions, insertions, and reversals. Doklady Akademii Nauk SSSR. 1965;163(4):845–848. (In Russ.).
2. Smith R. An overview of the Tesseract OCR engine. In: Proceedings of the Ninth International Conference on Document Analysis and Recognition (ICDAR 2007), 23–26 September 2007, Curitiba, Brazil. IEEE; 2007. P. 629–633. https://doi.org/10.1109/ICDAR.2007.4376991
3. Li Ch., Liu W., Guo R., et al. PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System. arXiv. URL: https://arxiv.org/abs/2206.03001 [Accessed 23rd March 2026].
4. Kureichik V.V., Rodzin S.I., Bova V.V. Deep learning methods for natural language text processing. Izvestiya SFedU. Engineering Sciences. 2022;(2):189–199. (In Russ.). https://doi.org/10.18522/2311-3103-2022-2-189-199
5. Shen Y., Wei J., Niu X., et al. Efficient ultra-lightweight convolutional attention network for embedded identity document recognition system. Image and Vision Computing. 2026;168:105930. https://doi.org/10.1016/j.imavis.2026.105930
6. Du Y., Li Ch., Guo R., et al. PP-OCR: A practical ultra lightweight OCR system. arXiv. URL: https://arxiv.org/abs/2009.09941 [Accessed 23rd March 2026].
7. Vaswani A., Shazeer N., Parmar N., et al. Attention Is All You Need. In: Advances in Neural Information Processing Systems 30 (NeurIPS 2017), 04–09 December 2017, Long Beach, CA, USA. 2017. P. 5998–6008.
8. Voloshchuk A.S. Metric books as a historical source: Characteristics, analysis, search and access to them. In: Document in Modern Society: Communicative Models and Technologies: Proceedings of the XVI All-Russian Student Scientific and Practical Conference, 07–08 April 2023, Yekaterinburg, Russia. Yekaterinburg: Ural Federal University; 2023. P. 256–260. (In Russ.).
9. Lvovich Ya.E., Khaknazarov I.V., Lyamzin I.S. On the possibilities of supporting electronic document management in organizations. In: Socio-economic development of Russia: problems, trends, prospects: Collection of scientific articles of participants of the 24th International Scientific and Practical Conference, 22 May 2025, Kursk, Russia. Kursk: Universitetskaya kniga; 2025. P. 299–301. (In Russ.).
10. Kuznetsov A.V. Beyond Topic Modeling: Analyzing Historical Text with Large Language Models. Historical Informatics. 2024;(4):47–65. (In Russ.). https://doi.org/10.7256/2585-7797.2024.4.72560
11. Bakulin A.Yu., Gusev P.Yu., Lvovich Ya.E. Optimizing the development of a digitalized organizational order servicing system based on integration with machine learning of predictive models. Vestnik of the Russian New University. Series: Complex Systems: models, analysis and management. 2026;(1):68–78. (In Russ.). https://doi.org/10.18137/RNU.V9187.26.01.P.68
Keywords: metric books, OCR, paddleOCR, gigaChat-MAX, LLM, pre-reform orthography, CER, WER, TF‑IDF
For citation: Kovaleva V.V., Gusev P.Y., Sepkin S.S. Development of an algorithm for recognizing archival texts based on the integration of OCR and large language models. Modeling, Optimization and Information Technology. 2026;14(8). URL: https://moitvivt.ru/ru/journal/article?id=2581 DOI: 10.26102/2310-6018/2026.59.8.002 (In Russ).
© Kovaleva V.V., Gusev P.Y., Sepkin S.S. Статья опубликована на условиях лицензии Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NS 4.0)Received 03.07.2026
Revised 31.07.2026
Accepted 10.08.2026