Keywords: information retrieval, e-commerce, keyboard layout correction, high-load systems, computational efficiency, technical search queries
UDC 005.94
DOI: 10.26102/2310-6018/2026.58.7.014
Modern e-commerce search systems operate under conditions of high user activity and the need to process short, structurally dense queries containing technical specifications and product identifiers. The "do-it-yourself" (DIY) segment presents particular challenges, as a substantial proportion of queries include alphanumeric tokens (e.g., "M8x50", "LED 12W", "SSD 1TB"). A significant source of retrieval quality degradation arises from keyboard layout errors (ENG→RU), in which Latin characters are interpreted as Cyrillic ones, producing unrecognized combinations (e.g., "ыыВ 1ЕИ" instead of "SSD 1TB"). Such distortions disrupt the lexical and semantic integrity of queries, reduce retrieval relevance, and negatively affect conversion rates. This study proposes a computationally efficient approach to ENG→RU keyboard layout correction designed for deployment in high-load information retrieval systems operating under limited computational resources. The method integrates deterministic key-mapping tables, contextual token matching, and lightweight machine learning models to achieve a balance between accuracy and performance. Experimental evaluation on domain-specific datasets demonstrates a 25–30 % improvement in retrieval accuracy compared to baseline spelling-correction methods, while maintaining an average response latency below 10 ms per query. The results confirm the scalability and industrial applicability of the proposed solution in server-side and cloud-based IR platforms.
1. Lee J.S., Choi K.S. English to Korean statistical transliteration for information retrieval. Comput Process Orient Lang. 1998;12(1):17–37.
2. Pogrebnoi D., Funkner A., Kovalchuk S. RuMedSpellchecker: correcting spelling errors for natural Russian language in electronic health records using machine learning techniques. In: International Conference on Computational Science, 3–5 July 2023, Prague, Czech Republic. Cham: Springer Nature Switzerland; 2023. p. 213–227. https://doi.org/10.1007/978-3-031-36027-5_17
3. Prabhakar D.K., Pal S. Machine transliteration and transliterated text retrieval: a survey. Sādhanā. 2018;43(6). https://doi.org/10.1007/s12046-018-0879-4
4. Balabaeva K., Funkner A., Kovalchuk S. Automated spelling correction for clinical text mining in Russian. In: Digital Personalized Health and Medicine, 17–19 November 2020, Amsterdam, Netherlands. Amsterdam: IOS Press; 2020. p. 43–47. https://doi.org/10.3233/SHTI200119
5. Rozovskaya A. Spelling correction for Russian: a comparative study of datasets and methods. In: Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021), 1–3 September 2021, Varna, Bulgaria. Shoumen: INCOMA Ltd.; 2021. p. 1206–1216.
6. Bruch S., Lucchese C., Maistro M., et al. Special section on efficiency in neural information retrieval. ACM Trans Inf Syst. 2024;42(5):1–4. https://doi.org/10.1145/3641203
7. Chari A., Ounis I., MacAvaney S. Lost in transliteration: bridging the script gap in neural IR. In: Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, 13–18 July 2025, Padua, Italy. New York: Association for Computing Machinery; 2025. p. 2900–2905.
8. Toutanova K., Moore R.C. Pronunciation modeling for improved spelling correction. In: Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 7–12 July 2002, Philadelphia, USA. Philadelphia: Association for Computational Linguistics; 2002. p. 144–151. https://doi.org/10.3115/1073083.1073110
9. Sachdeva N., McAuley J. How useful are reviews for recommendation? A critical review and potential improvements. In: Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, 25–30 July 2020, Xi'an, China. New York: Association for Computing Machinery; 2020. p. 1845–1848. https://doi.org/10.1145/3397271.3401286
10. Mikolov T., Chen K., Corrado G., et al. Efficient estimation of word representations in vector space. arXiv. URL: https://doi.org/10.48550/arXiv.1301.3781 [Accessed 1st February 2026].
11. Belinkov Y., Bisk Y. Synthetic and natural noise both break neural machine translation. arXiv. URL: https://doi.org/10.48550/arXiv.1711.02173 [Accessed 1st February 2026].
12. Müller L., Juozapavičius A., Okhrimchuk V., et al. Dictionary attack with transformed Russian words using QWERTY keyboard layout. Baltic Journal of Modern Computing. 2025;13(4):919–932. https://doi.org/10.31219/osf.io/mfqrw
13. Joulin A., Grave E., Bojanowski P., et al. Bag of tricks for efficient text classification. arXiv. URL: https://doi.org/10.48550/arXiv.1607.01759 [Accessed 1st February 2026].
14. Xue L., Constant N., Roberts A., et al. mT5: a massively multilingual pre-trained text-to-text transformer. arXiv. URL: https://doi.org/10.48550/arXiv.2010.11934 [Accessed 1st February 2026].
15. Lewis M., Liu Y., Goyal N., et al. BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv. URL: https://doi.org/10.48550/arXiv.1910.13461 [Accessed 1st February 2026].
16. Kudo T., Richardson J. SentencePiece: a simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv. URL: https://doi.org/10.48550/arXiv.1808.06226 [Accessed 1st February 2026].
Keywords: information retrieval, e-commerce, keyboard layout correction, high-load systems, computational efficiency, technical search queries
For citation: Krasnov F.V. Computationally efficient ENG→RU keyboard layout correction in high-load e-commerce information retrieval systems. Modeling, Optimization and Information Technology. 2026;14(7). URL: https://moitvivt.ru/ru/journal/article?id=2256 DOI: 10.26102/2310-6018/2026.58.7.014 (In Russ).
© Krasnov F.V. Статья опубликована на условиях лицензии Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NS 4.0)Received 27.02.2026
Revised 13.04.2026
Accepted 18.07.2026