Вычислительно эффективная коррекция раскладки ENG→RU в высоконагруженных информационно-поисковых системах электронной коммерции
Работая с сайтом, я даю свое согласие на использование файлов cookie. Это необходимо для нормального функционирования сайта, показа целевой рекламы и анализа трафика. Статистика использования сайта обрабатывается системой Яндекс.Метрика
Научный журнал Моделирование, оптимизация и информационные технологииThe scientific journal Modeling, Optimization and Information Technology
Online media
issn 2310-6018

Computationally efficient ENG→RU keyboard layout correction in high-load e-commerce information retrieval systems

idKrasnov F.V.

UDC 005.94
DOI: 10.26102/2310-6018/2026.58.7.014

  • Abstract
  • List of references
  • About authors

Modern e-commerce search systems operate under conditions of high user activity and the need to process short, structurally dense queries containing technical specifications and product identifiers. The "do-it-yourself" (DIY) segment presents particular challenges, as a substantial proportion of queries include alphanumeric tokens (e.g., "M8x50", "LED 12W", "SSD 1TB"). A significant source of retrieval quality degradation arises from keyboard layout errors (ENG→RU), in which Latin characters are interpreted as Cyrillic ones, producing unrecognized combinations (e.g., "ыыВ 1ЕИ" instead of "SSD 1TB"). Such distortions disrupt the lexical and semantic integrity of queries, reduce retrieval relevance, and negatively affect conversion rates. This study proposes a computationally efficient approach to ENG→RU keyboard layout correction designed for deployment in high-load information retrieval systems operating under limited computational resources. The method integrates deterministic key-mapping tables, contextual token matching, and lightweight machine learning models to achieve a balance between accuracy and performance. Experimental evaluation on domain-specific datasets demonstrates a 25–30 % improvement in retrieval accuracy compared to baseline spelling-correction methods, while maintaining an average response latency below 10 ms per query. The results confirm the scalability and industrial applicability of the proposed solution in server-side and cloud-based IR platforms.

1. Lee J.S., Choi K.S. English to Korean statistical transliteration for information retrieval. Comput Process Orient Lang. 1998;12(1):17–37.

2. Pogrebnoi D., Funkner A., Kovalchuk S. RuMedSpellchecker: correcting spelling errors for natural Russian language in electronic health records using machine learning techniques. In: International Conference on Computational Science, 3–5 July 2023, Prague, Czech Republic. Cham: Springer Nature Switzerland; 2023. p. 213–227. https://doi.org/10.1007/978-3-031-36027-5_17

3. Prabhakar D.K., Pal S. Machine transliteration and transliterated text retrieval: a survey. Sādhanā. 2018;43(6). https://doi.org/10.1007/s12046-018-0879-4

4. Balabaeva K., Funkner A., Kovalchuk S. Automated spelling correction for clinical text mining in Russian. In: Digital Personalized Health and Medicine, 17–19 November 2020, Amsterdam, Netherlands. Amsterdam: IOS Press; 2020. p. 43–47. https://doi.org/10.3233/SHTI200119

5. Rozovskaya A. Spelling correction for Russian: a comparative study of datasets and methods. In: Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021), 1–3 September 2021, Varna, Bulgaria. Shoumen: INCOMA Ltd.; 2021. p. 1206–1216.

6. Bruch S., Lucchese C., Maistro M., et al. Special section on efficiency in neural information retrieval. ACM Trans Inf Syst. 2024;42(5):1–4. https://doi.org/10.1145/3641203

7. Chari A., Ounis I., MacAvaney S. Lost in transliteration: bridging the script gap in neural IR. In: Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, 13–18 July 2025, Padua, Italy. New York: Association for Computing Machinery; 2025. p. 2900–2905.

8. Toutanova K., Moore R.C. Pronunciation modeling for improved spelling correction. In: Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 7–12 July 2002, Philadelphia, USA. Philadelphia: Association for Computational Linguistics; 2002. p. 144–151. https://doi.org/10.3115/1073083.1073110

9. Sachdeva N., McAuley J. How useful are reviews for recommendation? A critical review and potential improvements. In: Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, 25–30 July 2020, Xi'an, China. New York: Association for Computing Machinery; 2020. p. 1845–1848. https://doi.org/10.1145/3397271.3401286

10. Mikolov T., Chen K., Corrado G., et al. Efficient estimation of word representations in vector space. arXiv. URL: https://doi.org/10.48550/arXiv.1301.3781 [Accessed 1st February 2026].

11. Belinkov Y., Bisk Y. Synthetic and natural noise both break neural machine translation. arXiv. URL: https://doi.org/10.48550/arXiv.1711.02173 [Accessed 1st February 2026].

12. Müller L., Juozapavičius A., Okhrimchuk V., et al. Dictionary attack with transformed Russian words using QWERTY keyboard layout. Baltic Journal of Modern Computing. 2025;13(4):919–932. https://doi.org/10.31219/osf.io/mfqrw

13. Joulin A., Grave E., Bojanowski P., et al. Bag of tricks for efficient text classification. arXiv. URL: https://doi.org/10.48550/arXiv.1607.01759 [Accessed 1st February 2026].

14. Xue L., Constant N., Roberts A., et al. mT5: a massively multilingual pre-trained text-to-text transformer. arXiv. URL: https://doi.org/10.48550/arXiv.2010.11934 [Accessed 1st February 2026].

15. Lewis M., Liu Y., Goyal N., et al. BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv. URL: https://doi.org/10.48550/arXiv.1910.13461 [Accessed 1st February 2026].

16. Kudo T., Richardson J. SentencePiece: a simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv. URL: https://doi.org/10.48550/arXiv.1808.06226 [Accessed 1st February 2026].

Krasnov Fedor Vladimirovich
Candidate of Engineering Sciences

WoS | Scopus | ORCID | eLibrary |

Vi.Tech LLC

Kovrov, Russian Federation

Keywords: information retrieval, e-commerce, keyboard layout correction, high-load systems, computational efficiency, technical search queries

For citation: Krasnov F.V. Computationally efficient ENG→RU keyboard layout correction in high-load e-commerce information retrieval systems. Modeling, Optimization and Information Technology. 2026;14(7). URL: https://moitvivt.ru/ru/journal/article?id=2256 DOI: 10.26102/2310-6018/2026.58.7.014 (In Russ).

© Krasnov F.V. Статья опубликована на условиях лицензии Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NS 4.0)
18

Full text in PDF

Скачать JATS XML

Received 27.02.2026

Revised 13.04.2026

Accepted 18.07.2026