<?xml version="1.0" encoding="UTF-8"?>
<article article-type="research-article" dtd-version="1.3" xml:lang="ru" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:noNamespaceSchemaLocation="https://metafora.rcsi.science/xsd_files/journal3.xsd">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">moitvivt</journal-id>
      <journal-title-group>
        <journal-title xml:lang="ru">Моделирование, оптимизация и информационные технологии</journal-title>
        <trans-title-group xml:lang="en">
          <trans-title>Modeling, Optimization and Information Technology</trans-title>
        </trans-title-group>
      </journal-title-group>
      <issn pub-type="epub">2310-6018</issn>
      <publisher>
        <publisher-name>Издательство</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.26102/2310-6018/2026.59.8.002</article-id>
      <article-id pub-id-type="custom" custom-type="elpub">2581</article-id>
      <title-group>
        <article-title xml:lang="ru">Разработка алгоритма распознавания архивных текстов на основе интеграции OCR и больших языковых моделей</article-title>
        <trans-title-group xml:lang="en">
          <trans-title>Development of an algorithm for recognizing archival texts based on the integration of OCR and large language models</trans-title>
        </trans-title-group>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name-alternatives>
            <name name-style="eastern" xml:lang="ru">
              <surname>Ковалева</surname>
              <given-names>Виктория Викторовна</given-names>
            </name>
            <name name-style="western" xml:lang="en">
              <surname>Kovaleva</surname>
              <given-names>Victoria Viktorovna</given-names>
            </name>
          </name-alternatives>
          <email>vikakot0004@gmail.com</email>
          <xref ref-type="aff">aff-1</xref>
        </contrib>
        <contrib contrib-type="author">
          <contrib-id contrib-id-type="orcid">0000-0002-3752-0152</contrib-id>
          <name-alternatives>
            <name name-style="eastern" xml:lang="ru">
              <surname>Гусев</surname>
              <given-names>Павел Юрьевич</given-names>
            </name>
            <name name-style="western" xml:lang="en">
              <surname>Gusev</surname>
              <given-names>Pavel Yuryevich</given-names>
            </name>
          </name-alternatives>
          <email>gusevpvl@gmail.com</email>
          <xref ref-type="aff">aff-2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name-alternatives>
            <name name-style="eastern" xml:lang="ru">
              <surname>Сепкин</surname>
              <given-names>Сергей Сергеевич</given-names>
            </name>
            <name name-style="western" xml:lang="en">
              <surname>Sepkin</surname>
              <given-names>Sergei Sergeevich</given-names>
            </name>
          </name-alternatives>
          <email>sep001@gmail.com</email>
          <xref ref-type="aff">aff-3</xref>
        </contrib>
      </contrib-group>
      <aff-alternatives id="aff-1">
        <aff xml:lang="ru">Воронежский государственный технический университет</aff>
        <aff xml:lang="en">Voronezh State Technical University</aff>
      </aff-alternatives>
      <aff-alternatives id="aff-2">
        <aff xml:lang="ru">Воронежский государственный технический университет</aff>
        <aff xml:lang="en">Voronezh State Technical University</aff>
      </aff-alternatives>
      <aff-alternatives id="aff-3">
        <aff xml:lang="ru">Воронежский государственный технический университет</aff>
        <aff xml:lang="en">Voronezh State Technical University</aff>
      </aff-alternatives>
      <pub-date pub-type="epub">
        <day>01</day>
        <month>01</month>
        <year>2026</year>
      </pub-date>
      <volume>1</volume>
      <issue>1</issue>
      <elocation-id>10.26102/2310-6018/2026.59.8.002</elocation-id>
      <permissions>
        <copyright-statement>Copyright © Авторы, 2026</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/">
          <license-p>This work is licensed under a Creative Commons Attribution 4.0 International License</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://moitvivt.ru/ru/journal/article?id=2581"/>
      <abstract xml:lang="ru">
        <p>Статья посвящена повышению качества распознавания архивных текстов, представленных в виде метрических книг Российской империи конца XIX – начала XX века, являющихся уникальным историческим источником для генеалогии и демографии. Стандартные системы оптического распознавания символов, обученные на современных печатных текстах, демонстрируют на подобных документах точность не более 30 % из-за комплекса факторов: ветхость бумажных носителей, выцветание чернил, наличие дореформенной орфографии (буквы ѣ, і, ѳ, ѵ) и смешение кириллического и латинского начертаний. В работе проведён сравнительный анализ двух подходов к улучшению распознавания: дообучение модели PaddleOCR на исторических данных и использование в связке с большой языковой моделью GigaChat-MAX без дообучения. Для постобработки разработан специализированный запрос для языковой модели, включающий правила замены символов, словарь частотных терминов и правила восстановления дореформенной орфографии. Эксперименты на выборке из 50 разворотов метрической книги показали, что дообучение повышает точность до 48 %, а интеграция с языковой моделью – до 65 %, при этом коэффициент ошибок на символ снижается с 70 % до 35 %. Языковая модель исправляет более половины всех ошибок, особенно эффективно заменяя латиницу на кириллицу и восстанавливая исторические окончания. В результате исследования предложен алгоритм распознавания архивных рукописных документов, отличающийся применением комбинации большой языковой модели и модели OCR, и обеспечивающий снижение затрат на разметку данных и вычислительные ресурсы. Практическая значимость: предложенный подход позволяет автоматизировать распознавание метрических книг с точностью 65 %, сокращая ручной труд по транскрипции более чем вдвое.</p>
      </abstract>
      <trans-abstract xml:lang="en">
        <p>The paper addresses the problem of improving text recognition quality for metric books of the Russian Empire (late 19th – early 20th centuries), which are a valuable historical source for genealogy and demography. Conventional optical character recognition (OCR) systems trained on modern printed texts achieve an accuracy of no more than 30 % on such documents due to a combination of factors: paper deterioration, ink fading, pre-reform orthography (letters ѣ, і, ѳ, ѵ) and mixing of Cyrillic and Latin glyphs. The study presents a comparative analysis of two approaches: fine-tuning the PaddleOCR model on historical data versus using a pipeline of baseline PaddleOCR with the large language model GigaChat-MAX without fine-tuning. For post-processing, a specialised prompt was developed, including character substitution rules, a dictionary of frequent terms, and rules for restoring pre-reform spelling. Experiments on a sample of 50 spreads from a metric book showed that fine-tuning increases accuracy to 48 %, while the integration with the language model achieves 65 %, with the character error rate reduced from 70 % to 35 %. The language model corrects more than half of all errors, being particularly effective in replacing Latin letters with Cyrillic and restoring historical endings. The scientific novelty lies in the first direct comparison of the two strategies and the demonstration of the advantage of post-processing with a language model without the need for labelled data and computational costs. The practical significance is that the proposed approach enables automated recognition of metric books with 65 % accuracy, reducing manual transcription effort by more than half.</p>
      </trans-abstract>
      <kwd-group xml:lang="ru">
        <kwd>метрические книги</kwd>
        <kwd>OCR</kwd>
        <kwd>PaddleOCR</kwd>
        <kwd>GigaChat-MAX</kwd>
        <kwd>LLM</kwd>
        <kwd>дореформенная орфография</kwd>
        <kwd>CER</kwd>
        <kwd>WER</kwd>
        <kwd>TF‑IDF</kwd>
      </kwd-group>
      <kwd-group xml:lang="en">
        <kwd>metric books</kwd>
        <kwd>OCR</kwd>
        <kwd>PaddleOCR</kwd>
        <kwd>GigaChat-MAX</kwd>
        <kwd>LLM</kwd>
        <kwd>pre-reform orthography</kwd>
        <kwd>CER</kwd>
        <kwd>WER</kwd>
        <kwd>TF‑IDF</kwd>
      </kwd-group>
      <funding-group>
        <funding-statement xml:lang="ru">Исследование выполнено без спонсорской поддержки.</funding-statement>
        <funding-statement xml:lang="en">The study was performed without external funding.</funding-statement>
      </funding-group>
    </article-meta>
  </front>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="cit1">
        <label>1</label>
        <mixed-citation xml:lang="ru">Левенштейн В.И. Двоичные коды с исправлением выпадений, вставок и замещений символов. Доклады Академии наук СССР. 1965;163(4):845–848.</mixed-citation>
      </ref>
      <ref id="cit2">
        <label>2</label>
        <mixed-citation xml:lang="ru">Smith R. An overview of the Tesseract OCR engine. In: Proceedings of the Ninth International Conference on Document Analysis and Recognition (ICDAR 2007), 23–26 September 2007, Curitiba, Brazil. IEEE; 2007. P. 629–633. https://doi.org/10.1109/ICDAR.2007.4376991</mixed-citation>
      </ref>
      <ref id="cit3">
        <label>3</label>
        <mixed-citation xml:lang="ru">Li Ch., Liu W., Guo R., et al. PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System. arXiv. URL: https://arxiv.org/abs/2206.03001 [Accessed 23rd March 2026].</mixed-citation>
      </ref>
      <ref id="cit4">
        <label>4</label>
        <mixed-citation xml:lang="ru">Курейчик В.В., Родзин С.И., Бова В.В. Методы глубокого обучения для обработки текстов на естественном языке. Известия ЮФУ. Технические науки. 2022;(2):189–199. https://doi.org/10.18522/2311-3103-2022-2-189-199</mixed-citation>
      </ref>
      <ref id="cit5">
        <label>5</label>
        <mixed-citation xml:lang="ru">Shen Y., Wei J., Niu X., et al. Efficient ultra-lightweight convolutional attention network for embedded identity document recognition system. Image and Vision Computing. 2026;168:105930. https://doi.org/10.1016/j.imavis.2026.105930</mixed-citation>
      </ref>
      <ref id="cit6">
        <label>6</label>
        <mixed-citation xml:lang="ru">Du Y., Li Ch., Guo R., et al. PP-OCR: A practical ultra lightweight OCR system. arXiv. URL: https://arxiv.org/abs/2009.09941 [Accessed 23rd March 2026].</mixed-citation>
      </ref>
      <ref id="cit7">
        <label>7</label>
        <mixed-citation xml:lang="ru">Vaswani A., Shazeer N., Parmar N., et al. Attention Is All You Need. In: Advances in Neural Information Processing Systems 30 (NeurIPS 2017), 04–09 December 2017, Long Beach, CA, USA. 2017. P. 5998–6008.</mixed-citation>
      </ref>
      <ref id="cit8">
        <label>8</label>
        <mixed-citation xml:lang="ru">Волощук А.С. Метрические книги как исторический источник: характеристика, анализ, поиск и доступ к ним. В сборнике: Документ в современном обществе: коммуникативные модели и технологии: Материалы XVI Всероссийской студенческой научно-практической конференции, 07–08 апреля 2023 года, Екатеринбург, Россия. Екатеринбург: Уральский федеральный университет имени первого Президента России Б.Н. Ельцина; 2023. С. 256–260.</mixed-citation>
      </ref>
      <ref id="cit9">
        <label>9</label>
        <mixed-citation xml:lang="ru">Львович Я.Е., Хакназаров И.В., Лямзин И.С. О возможностях поддержки электронного документооборота в организациях. В сборнике: Социально-экономическое развитие России: проблемы, тенденции, перспективы: Сборник научных статей участников 24-й Международной научно-практической конференции, 22 мая 2025 года, Курск, Россия. Курск: Университетская книга; 2025. С. 299–301.</mixed-citation>
      </ref>
      <ref id="cit10">
        <label>10</label>
        <mixed-citation xml:lang="ru">Кузнецов А.В. За пределами тематического моделирования: анализ исторического текста с помощью больших языковых моделей. Историческая информатика. 2024;(4):47–65. https://doi.org/10.7256/2585-7797.2024.4.72560</mixed-citation>
      </ref>
      <ref id="cit11">
        <label>11</label>
        <mixed-citation xml:lang="ru">Бакулин А.Ю., Гусев П.Ю., Львович Я.Е. Оптимизация развития цифровизированной организационной системы обслуживания заказов на основе интеграции с машинным обучением прогностических моделей. Вестник Российского нового университета. Серия: Сложные системы: модели, анализ и управление. 2026;(1):68–78. https://doi.org/10.18137/RNU.V9187.26.01.P.68</mixed-citation>
      </ref>
    </ref-list>
    <fn-group>
      <fn fn-type="conflict">
        <p>The authors declare that there are no conflicts of interest present.</p>
      </fn>
    </fn-group>
  </back>
</article>