<?xml version="1.0" encoding="UTF-8"?>
<article article-type="research-article" dtd-version="1.3" xml:lang="ru" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:noNamespaceSchemaLocation="https://metafora.rcsi.science/xsd_files/journal3.xsd">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">moitvivt</journal-id>
      <journal-title-group>
        <journal-title xml:lang="ru">Моделирование, оптимизация и информационные технологии</journal-title>
        <trans-title-group xml:lang="en">
          <trans-title>Modeling, Optimization and Information Technology</trans-title>
        </trans-title-group>
      </journal-title-group>
      <issn pub-type="epub">2310-6018</issn>
      <publisher>
        <publisher-name>Издательство</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.26102/2310-6018/2026.60.9.005</article-id>
      <article-id pub-id-type="custom" custom-type="elpub">2447</article-id>
      <title-group>
        <article-title xml:lang="ru">Метод автоматического реферирования научных текстов на основе диагностически управляемой итерационной генерации</article-title>
        <trans-title-group xml:lang="en">
          <trans-title>A method for automatic summarization of scientific texts based on diagnostically driven iterative generation</trans-title>
        </trans-title-group>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <contrib-id contrib-id-type="orcid">0009-0002-4082-2925</contrib-id>
          <name-alternatives>
            <name name-style="eastern" xml:lang="ru">
              <surname>Гокарев</surname>
              <given-names>Вадим Николаевич</given-names>
            </name>
            <name name-style="western" xml:lang="en">
              <surname>Gokarev</surname>
              <given-names>Vadim Nikolaevich</given-names>
            </name>
          </name-alternatives>
          <email>vckain6387@gmail.com</email>
          <xref ref-type="aff">aff-1</xref>
        </contrib>
      </contrib-group>
      <aff-alternatives id="aff-1">
        <aff xml:lang="ru">Всероссийский институт научной и технической информации РАН</aff>
        <aff xml:lang="en">All-Russian Institute of Scientific and Technical Information of the Russian Academy of Sciences</aff>
      </aff-alternatives>
      <pub-date pub-type="epub">
        <day>01</day>
        <month>01</month>
        <year>2026</year>
      </pub-date>
      <volume>1</volume>
      <issue>1</issue>
      <elocation-id>10.26102/2310-6018/2026.60.9.005</elocation-id>
      <permissions>
        <copyright-statement>Copyright © Авторы, 2026</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/">
          <license-p>This work is licensed under a Creative Commons Attribution 4.0 International License</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://moitvivt.ru/ru/journal/article?id=2447"/>
      <abstract xml:lang="ru">
        <p>В работе рассматривается задача автоматического абстрактивного реферирования русскоязычных научных текстов с использованием локальных больших языковых моделей. Показано, что однопроходная генерация реферата языковой моделью без учета характеристик исходного документа приводит к систематическим типам ошибок: семантической неполноте, механическому копированию, информационной избыточности или недостаточности. Для выявления и устранения указанных недостатков предложена гибридная программная система, реализующая замкнутый контур управления процессом генерации с обратной связью от модуля трехфакторной диагностической оценки качества. Контур осуществляет итерационную коррекцию параметров запроса к языковой модели, таких как температура, количество подаваемых ключевых слов, формат инструкции системного запроса, минимальное и максимальное количество генерируемых токенов, на основе классификации диагностического профиля машинного реферата. В составе диагностической методики предложено асимметричное штрафное преобразование z-нормализованных отклонений, отражающее содержательное различие между избыточностью и недостаточностью лексической, семантической и компрессионной близости реферата к референсному распределению. Описана модульная архитектура системы, включающая модули итерационной генерации рефератов, диагностической оценки и блок автоматического останова итерационного процесса. Приведены результаты экспериментальной апробации на корпусе русскоязычных научных статей, продемонстрирован статистически значимый и устойчивый прирост качества по комбинированной диагностической метрике относительно однопроходной базовой линии на отложенной тестовой выборке.</p>
      </abstract>
      <trans-abstract xml:lang="en">
        <p>This paper examines the problem of automatic abstract summarization of Russian-language scientific texts using local large-scale language models. It is shown that single-pass abstract generation by a language model without taking into account the characteristics of the source document leads to systematic errors: semantic incompleteness, mechanical copying, and information redundancy or insufficiency. To identify and eliminate these deficiencies, a hybrid software system is proposed that implements a closed-loop control system for the abstract generation process with feedback from a three-factor diagnostic quality assessment module. The system iteratively adjusts query parameters to the language model, such as temperature, the number of submitted keywords, the format of the system query instruction, and the minimum and maximum number of generated tokens, based on the classification of the generated abstract diagnostic profile. The diagnostic methodology proposed includes an asymmetric penalty transformation of z-normalized deviations, reflecting the substantive difference between redundancies and insufficiencies in the lexical, semantic, and compressional proximity of an abstract to the reference distribution. The system's modular architecture is described, including modules for iterative abstract generation, diagnostic evaluation, and an automatic iteration process stop block. The results of experimental testing on a corpus of Russian-language scientific articles are presented, demonstrating a statistically significant and sustainable increase in quality for the combined diagnostic metric relative to a single-pass baseline on a delayed test sample.</p>
      </trans-abstract>
      <kwd-group xml:lang="ru">
        <kwd>автоматическое реферирование</kwd>
        <kwd>научные тексты</kwd>
        <kwd>большие языковые модели</kwd>
        <kwd>контур управления</kwd>
        <kwd>диагностическая оценка качества</kwd>
        <kwd>асимметричные штрафные функции</kwd>
      </kwd-group>
      <kwd-group xml:lang="en">
        <kwd>automatic summarization</kwd>
        <kwd>scientific texts</kwd>
        <kwd>large language models</kwd>
        <kwd>control loop</kwd>
        <kwd>diagnostic quality evaluation</kwd>
        <kwd>asymmetric penalty functions</kwd>
      </kwd-group>
      <funding-group>
        <funding-statement xml:lang="ru">Статья выполнена в рамках государственного задания: FFFU-2025-0010.</funding-statement>
        <funding-statement xml:lang="en">The article was prepared within the framework of the state assignment: FFFU-2025-0010.</funding-statement>
      </funding-group>
    </article-meta>
  </front>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="cit1">
        <label>1</label>
        <mixed-citation xml:lang="ru">Lewis M., Liu Y., Goyal N., et al. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 05–10 July 2020, Online. ACL; 2020. P. 7871–7880. https://doi.org/10.18653/v1/2020.acl-main.703</mixed-citation>
      </ref>
      <ref id="cit2">
        <label>2</label>
        <mixed-citation xml:lang="ru">Raffel C., Shazeer N., Roberts A., et al. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research. 2020;21(140):1–67.</mixed-citation>
      </ref>
      <ref id="cit3">
        <label>3</label>
        <mixed-citation xml:lang="ru">Gusev I. Dataset for Automatic Summarization of Russian News. In: Artificial Intelligence and Natural Language: 9th Conference (AINL 2020), 07–09 October 2020, Helsinki, Finland. Cham: Springer; 2020. P. 122–134. https://doi.org/10.1007/978-3-030-59082-6_9</mixed-citation>
      </ref>
      <ref id="cit4">
        <label>4</label>
        <mixed-citation xml:lang="ru">Balde G., Roy S., Mondal M., Ganguly N. Evaluation of LLMs in Medical Text Summarization: The Role of Vocabulary Adaptation in High OOV Settings. In: Findings of the Association for Computational Linguistics, 27 July – 01 August 2025, Vienna, Austria. ACL; 2025. P. 22989–23004. https://doi.org/10.18653/v1/2025.findings-acl.1179</mixed-citation>
      </ref>
      <ref id="cit5">
        <label>5</label>
        <mixed-citation xml:lang="ru">Keskar N.Sh., McCann B., Varshney L.R., Xiong C., Socher R. CTRL: A Conditional Transformer Language Model for Controllable Generation. arXiv. URL: https://arxiv.org/abs/1909.05858 [Accessed 6th March 2026].</mixed-citation>
      </ref>
      <ref id="cit6">
        <label>6</label>
        <mixed-citation xml:lang="ru">He J., Kryściński W., McCann B., Rajani N., Xiong C. CTRLsum: Towards Generic Controllable Text Summarization. arXiv. URL: https://arxiv.org/abs/2012.04281 [Accessed 6th March 2026].</mixed-citation>
      </ref>
      <ref id="cit7">
        <label>7</label>
        <mixed-citation xml:lang="ru">Lewis P., Perez E., Piktus A., et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In: Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020 (NeurIPS 2020), 06–12 December 2020, Online. 2020. URL: https://arxiv.org/abs/2005.11401</mixed-citation>
      </ref>
      <ref id="cit8">
        <label>8</label>
        <mixed-citation xml:lang="ru">Lin Ch.-Y. ROUGE: A Package for Automatic Evaluation of Summaries. In: Proceedings of Workshop on Text Summarization Branches Out: Post-Conference Workshop of ACL 2004, Barcelona, Spain. ACL; 2004. P. 74–81.</mixed-citation>
      </ref>
      <ref id="cit9">
        <label>9</label>
        <mixed-citation xml:lang="ru">Zhang T., Kishore V., Wu F., Weinberger K.Q., Artzi Y. BERTScore: Evaluating Text Generation with BERT. arXiv. URL: https://arxiv.org/abs/1904.09675 [Accessed 6th March 2026].</mixed-citation>
      </ref>
      <ref id="cit10">
        <label>10</label>
        <mixed-citation xml:lang="ru">Sellam Th., Das D., Parikh A.P. BLEURT: Learning Robust Metrics for Text Generation. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 05–10 July 2020, Online. ACL; 2020. P. 7881–7892. https://doi.org/10.18653/v1/2020.acl-main.704</mixed-citation>
      </ref>
      <ref id="cit11">
        <label>11</label>
        <mixed-citation xml:lang="ru">Maynez J., Narayan Sh., Bohnet B., McDonald R.T. On Faithfulness and Factuality in Abstractive Summarization. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 05–10 July 2020, Online. ACL; 2020. P. 1906–1919. https://doi.org/10.18653/v1/2020.acl-main.173</mixed-citation>
      </ref>
      <ref id="cit12">
        <label>12</label>
        <mixed-citation xml:lang="ru">Vasilyev O.V., Dharnidharka V., Bohannon J. Fill in the BLANC: Human-free Quality Estimation of Document Summaries. In: Proceedings of the First Workshop on Evaluation and Comparison of NLP Systems (Eval4NLP 2020), 20 November 2020, Online. ACL; 2020. P. 11–20. https://doi.org/10.18653/v1/2020.eval4nlp-1.2</mixed-citation>
      </ref>
      <ref id="cit13">
        <label>13</label>
        <mixed-citation xml:lang="ru">Rose S., Engel D., Cramer N., Cowley W. Automatic Keyword Extraction from Individual Documents. In: Text Mining: Applications and Theory. Chichester: John Wiley &amp; Sons; 2010. P. 1–20. https://doi.org/10.1002/9780470689646.ch1</mixed-citation>
      </ref>
      <ref id="cit14">
        <label>14</label>
        <mixed-citation xml:lang="ru">Mihalcea R., Tarau P. TextRank: Bringing Order into Text. In: Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing (EMNLP 2004), 25–26 July 2004, Barcelona, Spain. ACL; 2004. P. 404–411.</mixed-citation>
      </ref>
      <ref id="cit15">
        <label>15</label>
        <mixed-citation xml:lang="ru">Лукашевич Н.В., Логачёв Ю.М. Комбинирование признаков для автоматического извлечения терминов. Вычислительные методы и программирование. 2010;11(4):108–116.</mixed-citation>
      </ref>
      <ref id="cit16">
        <label>16</label>
        <mixed-citation xml:lang="ru">Liu N.F., Lin K., Hewitt J., et al. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics. 2024;12:157–173. https://doi.org/10.1162/tacl_a_00638</mixed-citation>
      </ref>
      <ref id="cit17">
        <label>17</label>
        <mixed-citation xml:lang="ru">Hsieh Ch.-Y., Chuang Y.-S., Li Ch.-L., et al. Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization. In: Findings of the Association for Computational Linguistics (ACL 2024), 11–16 August 2024, Bangkok, Thailand, Online. ACL; 2024. P. 14982–14995. https://doi.org/10.18653/v1/2024.findings-acl.890</mixed-citation>
      </ref>
      <ref id="cit18">
        <label>18</label>
        <mixed-citation xml:lang="ru">Xu L., Karim M.A., Dingliwal S., Elangovan A. Salient Information Prompting to Steer Content in Prompt-based Abstractive Summarization. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP 2024): Industry Track, 12–16 November 2024, Miami, FL, USA. ACL; 2024. P. 35–49. https://doi.org/10.18653/v1/2024.emnlp-industry.4</mixed-citation>
      </ref>
    </ref-list>
    <fn-group>
      <fn fn-type="conflict">
        <p>The authors declare that there are no conflicts of interest present.</p>
      </fn>
    </fn-group>
  </back>
</article>