Метод автоматического реферирования научных текстов на основе диагностически управляемой итерационной генерации
Работая с сайтом, я даю свое согласие на использование файлов cookie. Это необходимо для нормального функционирования сайта, показа целевой рекламы и анализа трафика. Статистика использования сайта обрабатывается системой Яндекс.Метрика
Научный журнал Моделирование, оптимизация и информационные технологииThe scientific journal Modeling, Optimization and Information Technology
Online media
issn 2310-6018

A method for automatic summarization of scientific texts based on diagnostically driven iterative generation

idGokarev V.N.

UDC 004.912
DOI: 10.26102/2310-6018/2026.60.9.005

  • Abstract
  • List of references
  • About authors

This paper examines the problem of automatic abstract summarization of Russian-language scientific texts using local large-scale language models. It is shown that single-pass abstract generation by a language model without taking into account the characteristics of the source document leads to systematic errors: semantic incompleteness, mechanical copying, and information redundancy or insufficiency. To identify and eliminate these deficiencies, a hybrid software system is proposed that implements a closed-loop control system for the abstract generation process with feedback from a three-factor diagnostic quality assessment module. The system iteratively adjusts query parameters to the language model, such as temperature, the number of submitted keywords, the format of the system query instruction, and the minimum and maximum number of generated tokens, based on the classification of the generated abstract diagnostic profile. The diagnostic methodology proposed includes an asymmetric penalty transformation of z-normalized deviations, reflecting the substantive difference between redundancies and insufficiencies in the lexical, semantic, and compressional proximity of an abstract to the reference distribution. The system's modular architecture is described, including modules for iterative abstract generation, diagnostic evaluation, and an automatic iteration process stop block. The results of experimental testing on a corpus of Russian-language scientific articles are presented, demonstrating a statistically significant and sustainable increase in quality for the combined diagnostic metric relative to a single-pass baseline on a delayed test sample.

1. Lewis M., Liu Y., Goyal N., et al. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 05–10 July 2020, Online. ACL; 2020. P. 7871–7880. https://doi.org/10.18653/v1/2020.acl-main.703

2. Raffel C., Shazeer N., Roberts A., et al. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research. 2020;21(140):1–67.

3. Gusev I. Dataset for Automatic Summarization of Russian News. In: Artificial Intelligence and Natural Language: 9th Conference (AINL 2020), 07–09 October 2020, Helsinki, Finland. Cham: Springer; 2020. P. 122–134. https://doi.org/10.1007/978-3-030-59082-6_9

4. Balde G., Roy S., Mondal M., Ganguly N. Evaluation of LLMs in Medical Text Summarization: The Role of Vocabulary Adaptation in High OOV Settings. In: Findings of the Association for Computational Linguistics, 27 July – 01 August 2025, Vienna, Austria. ACL; 2025. P. 22989–23004. https://doi.org/10.18653/v1/2025.findings-acl.1179

5. Keskar N.Sh., McCann B., Varshney L.R., Xiong C., Socher R. CTRL: A Conditional Transformer Language Model for Controllable Generation. arXiv. URL: https://arxiv.org/abs/1909.05858 [Accessed 6th March 2026].

6. He J., Kryściński W., McCann B., Rajani N., Xiong C. CTRLsum: Towards Generic Controllable Text Summarization. arXiv. URL: https://arxiv.org/abs/2012.04281 [Accessed 6th March 2026].

7. Lewis P., Perez E., Piktus A., et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In: Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020 (NeurIPS 2020), 06–12 December 2020, Online. 2020. URL: https://arxiv.org/abs/2005.11401

8. Lin Ch.-Y. ROUGE: A Package for Automatic Evaluation of Summaries. In: Proceedings of Workshop on Text Summarization Branches Out: Post-Conference Workshop of ACL 2004, Barcelona, Spain. ACL; 2004. P. 74–81.

9. Zhang T., Kishore V., Wu F., Weinberger K.Q., Artzi Y. BERTScore: Evaluating Text Generation with BERT. arXiv. URL: https://arxiv.org/abs/1904.09675 [Accessed 6th March 2026].

10. Sellam Th., Das D., Parikh A.P. BLEURT: Learning Robust Metrics for Text Generation. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 05–10 July 2020, Online. ACL; 2020. P. 7881–7892. https://doi.org/10.18653/v1/2020.acl-main.704

11. Maynez J., Narayan Sh., Bohnet B., McDonald R.T. On Faithfulness and Factuality in Abstractive Summarization. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 05–10 July 2020, Online. ACL; 2020. P. 1906–1919. https://doi.org/10.18653/v1/2020.acl-main.173

12. Vasilyev O.V., Dharnidharka V., Bohannon J. Fill in the BLANC: Human-free Quality Estimation of Document Summaries. In: Proceedings of the First Workshop on Evaluation and Comparison of NLP Systems (Eval4NLP 2020), 20 November 2020, Online. ACL; 2020. P. 11–20. https://doi.org/10.18653/v1/2020.eval4nlp-1.2

13. Rose S., Engel D., Cramer N., Cowley W. Automatic Keyword Extraction from Individual Documents. In: Text Mining: Applications and Theory. Chichester: John Wiley & Sons; 2010. P. 1–20. https://doi.org/10.1002/9780470689646.ch1

14. Mihalcea R., Tarau P. TextRank: Bringing Order into Text. In: Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing (EMNLP 2004), 25–26 July 2004, Barcelona, Spain. ACL; 2004. P. 404–411.

15. Lukashevich N.V., Logachev Yu.M. Automatic term extraction based on feature combination. Numerical Methods and Programming. 2010;11(4):108–116. (In Russ.).

16. Liu N.F., Lin K., Hewitt J., et al. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics. 2024;12:157–173. https://doi.org/10.1162/tacl_a_00638

17. Hsieh Ch.-Y., Chuang Y.-S., Li Ch.-L., et al. Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization. In: Findings of the Association for Computational Linguistics (ACL 2024), 11–16 August 2024, Bangkok, Thailand, Online. ACL; 2024. P. 14982–14995. https://doi.org/10.18653/v1/2024.findings-acl.890

18. Xu L., Karim M.A., Dingliwal S., Elangovan A. Salient Information Prompting to Steer Content in Prompt-based Abstractive Summarization. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP 2024): Industry Track, 12–16 November 2024, Miami, FL, USA. ACL; 2024. P. 35–49. https://doi.org/10.18653/v1/2024.emnlp-industry.4

Gokarev Vadim Nikolaevich

ORCID |

All-Russian Institute of Scientific and Technical Information of the Russian Academy of Sciences

Moscow, Russian Federation

Keywords: automatic summarization, scientific texts, large language models, control loop, diagnostic quality evaluation, asymmetric penalty functions

Sources of funding: The article was prepared within the framework of the state assignment: FFFU-2025-0010.

For citation: Gokarev V.N. A method for automatic summarization of scientific texts based on diagnostically driven iterative generation. Modeling, Optimization and Information Technology. 2026;14(9). URL: https://moitvivt.ru/ru/journal/article?id=2447 DOI: 10.26102/2310-6018/2026.60.9.005 (In Russ).

© Gokarev V.N. Статья опубликована на условиях лицензии Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NS 4.0)
14

Full text in PDF

Скачать JATS XML

Received 23.05.2026

Revised 09.09.2026

Accepted 15.09.2026