HEИндивидуална стипендия2025–2027

RAVIOLI · Retrieval-Augmented VIsion-Language Models for Open-vocabulary LocalizatIon

„Хоризонт Европа“ — Действия „Мария Склодовска-Кюри“

Период
2025-09-01 → 2027-08-31
Финансиране от ЕС
191 918 €
Участници
2
Схема
HORIZON-TMA-MSCA-PF-EF

Линиите свързват координатора с партньорите.

Накратко на български

Визуално-езиковите модели се подобряват чрез добавяне на външна памет, за да разпознават по-точно обекти в изображения, например в медицински снимки или при автономни автомобили. Това помага на системите по-лесно да се адаптират към нови и сложни видове обекти.

Този кратък обзор е генериран от изкуствен интелект

Кратко обяснение, генерирано от езиков модел по текста на CORDIS. Оригиналът е по-долу.

Цел на проекта

The proposed research project, RAVIOLI (Retrieval-Augmented VIsion-Language Models for Open-vocabulary LocalizatIon), aims to significantly advance the field of segmentation by innovatively integrating retrieval-based predictions from a memory with the original predictions of a vision-language model (VLM) through a learnable fusion model. Addressing a critical gap in existing methods, which often struggle to adapt to new or complex classes and domains, RAVIOLI seeks to enhance the accuracy, adaptability, and granularity of segmentation tasks across various applications, from autonomous vehicles to medical imaging. Importantly, there has been no similar attempt to learn a fusion model with these properties in any open-vocabulary dense task, such as segmentation, making our approach truly pioneering. The ambitious scope of this project lies in its aim to create a tailored, flexible, robust, and scalable solution that will redefine the capabilities of vision-language models, setting a new standard in the field of open-vocabulary segmentation. The project will be hosted by the Visual Recognition Group (VRG) at the Czech Technical University in Prague (CTU) under the supervision of Prof. Giorgos Tolias. The fellow, Bill Psomas, with a strong background in computer vision (CV) and deep learning (DL), is well-equipped to lead this research, which will further supported by a secondment at AImageLab, University of Modena and Reggio Emilia (UNIMORE) working with Prof. Rita Cucchiara.

Оригинален текст от CORDIS (на английски).

Участници

  • CESKE VYSOKE UCENI TECHNICKE V PRAZE · PRAHAКоординаторЧехия
  • UNIVERSITA DEGLI STUDI DI MODENA E REGGIO EMILIA · ModenaИталия

Връзки

Данни: CORDIS, © Европейски съюз