Определение геолокации по кадрам видеорегистратора с использованием ResNet
Работая с сайтом, я даю свое согласие на использование файлов cookie. Это необходимо для нормального функционирования сайта, показа целевой рекламы и анализа трафика. Статистика использования сайта обрабатывается системой Яндекс.Метрика
Научный журнал Моделирование, оптимизация и информационные технологииThe scientific journal Modeling, Optimization and Information Technology
Online media
issn 2310-6018

Geolocation determination from dashcam footage using ResNet

idBaskhanov A.R.

UDC 004.93
DOI: 10.26102/2310-6018/2026.58.7.016

  • Abstract
  • List of references
  • About authors

The relevance of this study is driven by the growing volume of street‑level imagery and the need for automatic georeferencing without relying on GPS metadata. Accordingly, this paper aims to assess the feasibility of determining photograph coordinates within a single city using deep neural networks. The leading research method is the application of convolutional neural networks ResNet18 and ResNet50 for latitude and longitude regression from 512×512-pixel frames. Data were collected using the open Mapillary platform, and a custom program with a 10,000‑cell grid was developed to ensure uniform coverage of Ufa. A total of 88,798 geotagged images were collected. The paper presents quantitative results of model training. The ResNet50‑based model achieves a median localization error of approximately 3.1 km after 13 epochs, with 92.6% of predictions falling within a 10 km radius of the ground truth. Smooth Grad‑CAM visualizations reveal that the model focuses on building facades, intersections, road markings, and other infrastructure elements – semantically meaningful landmarks for urban geolocation. The materials of the article are of practical value for automatic georeferencing systems for large collections of street‑level imagery and for the development of visual localization methods for autonomous vehicles. The work also provides a reproducible pipeline for city‑level image geolocation and analyses the behavioral attention of convolutional neural networks in this task.

1. Weyand T., Kostrikov I., Philbin J. PlaNet – Photo Geolocation with Convolutional Neural Networks. In: Proceedings of the Computer Vision – ECCV 2016: 14th European Conference, 11–14 October 2016, Amsterdam, Netherlands. Cham: Springer; 2016. P. 37–55. https://doi.org/10.1007/978-3-319-46484-8_3

2. Hays J., Efros A.A. IM2GPS: estimating geographic information from a single image. In: Proceedings of the 2008 IEEE Conference on Computer Vision and Pattern Recognition, 23–28 June 2008, Anchorage, AK, USA. IEEE; 2008. P. 1–8. https://doi.org/10.1109/CVPR.2008.4587784

3. Workman S., Jacobs N. On the location dependence of convolutional neural network features. In: Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 7–12 June 2015, Boston, MA, USA. IEEE; 2015. P. 70–78. https://doi.org/10.1109/CVPRW.2015.7301385

4. Melekhov I., Kannala J., Rahtu E. Siamese network features for image matching. In: Proceedings of the 2016 23rd International Conference on Pattern Recognition (ICPR), 4–8 December 2016, Cancún, Mexico. IEEE; 2016. P. 378–383. https://doi.org/10.1109/ICPR.2016.7899663

5. Hou Y., Quintana M., Khomiakov M., et al. Global Streetscapes – a comprehensive dataset of 10 million street-level images across 688 cities for urban science and analytics. ISPRS Journal of Photogrammetry and Remote Sensing. 2024;215:216–238. https://doi.org/10.1016/j.isprsjprs.2024.06.023

6. Chattopadhyay A., Sarkar A., Howlader P., et al. Grad-CAM++: Generalized Gradient-Based Visual Explanations for Deep Convolutional Networks. In: Proceedings of the 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), 12–15 March 2018, Lake Tahoe, NV, USA. IEEE; 2018. P. 839–847. https://doi.org/10.1109/WACV.2018.00097

7. He K., Zhang X., Ren S., et al. Deep Residual Learning for Image Recognition. In: Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 27–30 June 2016, Las Vegas, NV, USA. IEEE; 2016. P. 770–778. https://doi.org/10.1109/CVPR.2016.90

8. Wu Z., Shen C., Van Den Hengel A. Wider or Deeper: Revisiting the ResNet Model for Visual Recognition. Pattern Recognition. 2019;90:119–133. https://doi.org/10.1016/j.patcog.2019.01.006

9. Selvaraju R.R., Cogswell M., Das A., et al. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. International Journal of Computer Vision. 2020;128(2):336–359. https://doi.org/10.1007/s11263-019-01228-7

10. Shimizu T., Nagata F., Arima K., et al. Enhancing defective region visualization in industrial products using Grad-CAM and random masking data augmentation. Artificial Life and Robotics. 2024;29:62–69. https://doi.org/10.1007/s10015-023-00913-8

11. Müller-Budack E., Pustu-Iren K., Ewerth R. Geolocation Estimation of Photos Using a Hierarchical Model and Scene Classification. In: Proceedings of the 15th European Conference, 8–14 September 2018, Munich, Germany. Cham: Springer; 2018. P. 575–592. https://doi.org/10.1007/978-3-030-01258-8_35

Baskhanov Artur Ruslanovich

ORCID |

Financial University under the Government of the Russian Federation

Moscow, Russian Federation

Keywords: image geolocation, convolutional neural networks, resNet, urban scene analysis, deep regression, attention visualization, smooth Grad‑CAM, mapillary

For citation: Baskhanov A.R. Geolocation determination from dashcam footage using ResNet. Modeling, Optimization and Information Technology. 2026;14(7). URL: https://moitvivt.ru/ru/journal/article?id=2399 DOI: 10.26102/2310-6018/2026.58.7.016 (In Russ).

© Baskhanov A.R. Статья опубликована на условиях лицензии Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NS 4.0)
23

Full text in PDF

Скачать JATS XML

Received 03.05.2026

Revised 26.06.2026

Accepted 17.07.2026