VUAD · Video Understanding for Autonomous Driving
„Хоризонт 2020“ — Действия „Мария Склодовска-Кюри“
- Период
- 2020-04-01 → 2022-03-31
- Финансиране от ЕС
- 145 356 €
- Участници
- 1
- Схема
- MSCA-IF-EF-ST
Линиите свързват координатора с партньорите.
Накратко на български
Алгоритмите за автономно шофиране се подобряват чрез анализ на видео последователности, вместо само на единични снимки, за да се разпознава по-добре непрекъснатостта на обектите. Това повишава надеждността на машините и помага за намаляване на катастрофите, причинени от човешка грешка.
Кратко обяснение, генерирано от езиков модел по текста на CORDIS. Оригиналът е по-долу.
Резултати накратко
Video Understanding for Autonomous Driving
In this project, we address scene understanding from video sequences for autonomous vehicles. With the development of robust techniques in deep learning, the dream of autonomous vehicles has become a vision. However, the perception module for autonomous vehicles needs to be perfected to a level comparable to humans before these machines can be safely utilized on the roads. The goal of this project is to increase the robustness and reliability of perception algorithms in autonomous vehicles by modeling temporal cues in a video such as continuity as opposed to most state-of-the-art methods that use only a single image. Self-driving cars are poised to become a trillion-pound market in the next few decades, based on the needs of commuters and logistics chains worldwide. More importantly, they would solve two pressing problems of our society. 1.25 million people die in car accidents each year due to human error, and about 35 million are severely injured, rivaling the worst diseases. Another often-ignored fact is that the average car commuter spends 52 minutes per day driving to or from work, amounting to 5.4% of their waking time lost to a menial task. Enabling an Artificial Intelligence (AI) system to understand and drive in complex urban environments now seems largely solved for most common scenarios. Companies such as Waymo (US), Tesla (US), and Wayve (UK) routinely test on public roads. This achievement was made possible, largely, by advances in deep neural networks. Computer Vision methods achieve impressive results on a single image for various tasks such as object detection. For instance, pedestrian detectors now boast over 98% accuracy according to the widely-acknowledged KITTI benchmark. However, this success has not been fully extended to sequences yet. It is commonly acknowledged that video understanding falls years behind a single image. This is mainly due to two reasons: the processing power required for reasoning across multiple frames and the difficulty of obtaining ground truth for every frame in a sequence, especially for pixel-level tasks. Based on these observations, there are two likely directions to boost the performance of tasks related to video understanding: unsupervised learning and object-level reasoning. We work on both perspectives in this project. We present deep learning solutions for dynamic scene understanding by detecting and tracking multiple people in street scenes, i.e. multi-object tracking (MOT) as well as by modeling the movement of the static parts of the scene which arise from camera motion.
Текст от CORDIS, на английски · Данни: CORDIS, © Европейски съюз
Цел на проекта
Autonomous vision aims to solve computer vision problems related to autonomous driving. Autonomous vision algorithms achieve impressive results on a single image for various tasks such as object detection and semantic segmentation, however, this success has not been fully extended to video sequences yet. In computer vision, it is commonly acknowledged that video understanding falls years behind single image. This is mainly due to two reasons: processing power required for reasoning across multiple frames and the difficulty of obtaining ground truth for every frame in a sequence, especially for pixel-level tasks such as motion estimation. Based on these observations, there are two likely directions to boost the performance of tasks related to video understanding in autonomous vision: unsupervised learning and object-level reasoning as opposed to pixel-level reasoning. Following these directions, we propose to tackle three relevant problems in video understanding. First, we propose a deep learning method for multi-object tracking on graph structured data. Second, we extend it to joint video object detection and tracking by exploiting temporal cues in order to improve both detection and tracking performance. Third, we propose to learn a background motion model for the static parts of the scene in an unsupervised manner. Our long-term goal is also to be able to learn detection and tracking in an unsupervised manner. Once we achieve these stepping stones, we plan to combine the proposed algorithms into a unified video understanding module and test its performance in comparison to static counterparts as well as the state-of-the-art algorithms in video understanding.
Оригинален текст от CORDIS (на английски).
Участници
- KOC UNIVERSITY · IstanbulКоординаторТурция
Връзки
Данни: CORDIS, © Европейски съюз
