RoSTBiDFramework · Optimised Framework based on Rough Set Theory for Big Data Pre-processing in Certain and Imprecise Contexts
„Хоризонт 2020“ — Действия „Мария Склодовска-Кюри“
- Период
- 2017-03-01 → 2019-02-28
- Финансиране от ЕС
- 183 455 €
- Участници
- 2
- Схема
- MSCA-IF-EF-ST
Линиите свързват координатора с партньорите.
Накратко на български
Разработва се нов софтуерен модел за почистване на огромни масиви от данни, които съдържат грешки или липсваща информация. Това помага на бизнеса да извлича по-точни анализи и да взема по-добри решения за своите продукти и услуги.
Кратко обяснение, генерирано от езиков модел по текста на CORDIS. Оригиналът е по-долу.
Резултати накратко
Optimised Framework based on Rough Set Theory for Big Data Pre-processing in Certain and Imprecise Contexts
Big data is characterized by its Volume, Variety, Velocity, and Veracity. These 4Vs present complex limitations for enabling the derivation of meaningful information from big data. Hence, fixing the 4Vs is crucial to get valuable data. However, when dealing with these 4Vs, standard techniques have several limitations: (1) They depend on expert knowledge, which could mislead decision making with their possible subjective advices. (2) They are unable to decide how trustful data is when encountering missing data. (3) They are unable to deal with the intensive big data computations. Thus, these techniques are ineffective and they cannot fix the 4Vs at once. This project’s overarching aim is to fill these research gaps by developing an optimized framework, dubbed RoSTBiDFramework, for big data pre-processing in certain and imprecise contexts. RoSTBiDFramework comprises mathematical models, namely Rough Set Theory (RST), for data pre-processing, hashing techniques for optimization, and distributed frameworks. This innovative idea of hybridizing different disciplines is the key to fix the 4Vs to help decision makers. By applying RoSTBiDFramework to large pools of data gathered from any discipline/sector, only valuable and fine-grained data is generated to discern patterns and improve decision-making. The RoSTBiDFramework output, i.e., the pre-processed data, will become the main source of competition and growth for a variety of industries and businesses, allowing them to enhance their productivity and to create important value for the world economy. This is by increasing the quality of products and services, by minimizing risks, and by unearthing valuable insights that would otherwise remain hidden. In such a way, RoSTBiDFramework will help companies to have much more solid basics to outperform their peers. RoSTBiDFramework is based on five interconnected Research Objectives (RO) (RO1) The design and implementation of a feature selection framework based on RST in a certain context (RO2) The formalisation and implementation of an optimised version of the framework (RO3) The derivation of a general formulation of RST to deal with missing attribute values (RO4) The formalisation and implementation of an optimised version of the framework for handling the veracity aspect (RO5) The demonstration of RoSTBiDFramework on real-world data.
Текст от CORDIS, на английски · Данни: CORDIS, © Европейски съюз
Цел на проекта
Over the last decades, the amount of data has increased in an unprecedented rate, leading to a new terminology: ""Big Data"". Big data are specified by their Volume, Variety, Velocity and by their Veracity/Imprecision. Based on these 4V specificities, it has become difficult to quickly acquire the most useful information from the huge amount of data at hand. Thus, it is necessary to perform data (pre-)processing as a first step. In spite of the existence of many techniques for this task, most of the state-of-the-art methods require additional information for thresholding and are neither able to deal with the big data veracity aspect nor with their computational requirements. This project's overarching aim is to fill these major research gaps with an optimised framework for big data pre-processing in certain and imprecise contexts. Our approach is based on Rough Set Theory (RST) for data pre-processing and Randomised Search Heuristics for optimisation and will be implemented under the Spark MapReduce model.The project combines the expertise of the experienced researcher Dr Zaineb Chelly Dagdia in machine learning, rough set theory and information extraction with the knowledge in optimisation and randomised search heuristics of the supervisor Dr Christine Zarges at the University of Birmingham (UoB). Further expertise is provided by internal and external collaborators from academic and non-academic institutions, namely Prof Tino (UoB), Prof Merelo (University of Granada), Prof Lebbah (University of Paris 13) and Philippe Barra (Arrow Group). The involvement of Arrow Group, an SME based in France specialised in Big data, Banking, Finance & Insurance is of particular importance to ensure that real-world requirements are met throughout the development of the framework.""
Оригинален текст от CORDIS (на английски).
Участници
- ABERYSTWYTH UNIVERSITY · AberystwythКоординаторОбединеното кралство
- THE UNIVERSITY OF BIRMINGHAM · BirminghamОбединеното кралство
Връзки
Данни: CORDIS, © Европейски съюз
