FP7Реинтеграция2010–2014

EVALUATE · Theory and Practice of Algorithms for analysis of People and Data on the Web

7РП — „Хора“ (Действия „Мария Кюри“)

Период
2010-09-01 → 2014-08-31
Финансиране от ЕС
100 000 €
Участници
1
Схема
MC-IRG

Линиите свързват координатора с партньорите.

Накратко на български

Алгоритмите за анализ на данни проучват как да се подреждат предпочитанията на хората в мрежата и как да се групират данни чрез математически трансформации. Това помага за по-доброто разбиране на теоретичните основи, върху които се изграждат системите за работа с големи масиви от данни.

Този кратък обзор е генериран от изкуствен интелект

Кратко обяснение, генерирано от езиков модел по текста на CORDIS. Оригиналът е по-долу.

Резултати накратко

Theory and Practice of Algorithms for analysis of People and Data on the Web

The project set out to study 3 main scientific goals, namely: (1) Understanding learning theoretical and social choice theoretical aspects of learning to rank from preferences, (2) Studying problems related to the Correlation Clustering problem with prior trust information and (3) improve our understanding of the FJLT (Fast Johnson-Lindenstrauss Transforms). The expected impact of this work was to better understand the theoretical aspects underpinning some of the most important building blocks of the world of big data which is increasingly rapidly affecting many aspects of the lives of billions. For objective (1), the PI and his team have made progress in several directions. In [1,2], new active learning algorithms over pairwise preferences were designed, giving rise to new techniques for active learning. The new techniques are called “Smooth Relative Regret Approximation” (SRRA), and they are interesting in their own right in learning theory. In [7,9] the PI and coauthors were able to design online learning algorithms for an algorithm playing over an action set consisting of all permutations, in an environment incurring a loss function that is linear in a standard representation of the permutation. The algorithms bridged between almost optimal regret bounds as well as computational bounds. In [8] the PI and his collaborators defined a natural extension of a problem called “bandit stochastic optimization” to a pairwise case, in which the binary feedback from choosing a pair of actions is based on the outcome of a duel between the actions. For objective (2), the progress was made in the following aspects. The new techniques for active learning mentioned above (cf. [1]) turned out to be useful also in devising learning algorithms for a setting in which similarity labels are provided for pairs of objects. The labels are noisy, and possibly adversarily. In [5], a technique known as trace-norm minimization was used to solve a clustering problem known as “planted partitioning” (which is a stochastic case of correlation clustering). Unlike previous results, our result used a notion of adaptivity. In [6], we have developed a new analysis of an algorithm knows as k-means++ for a problem related to correlation clustering, and showed that the problem is more difficult that previously assumed. For objective (3), we have shown in [4] a strong connection between FJLT and so-called RIP matrices (Restricted Isometry Property) by defining a carefully constructed random walk in the unitary group. Studying the properties of this random walk, defined by repeatedly randomly flipping signs and applying a Fourier transform, opened the door to further open problems. In [10] I have shown new lower bounds for Fourier transform computation, which is a core step in FJLT.

Текст от CORDIS, на английски · Данни: CORDIS, © Европейски съюз

Цел на проекта

Since the late 20th century, an ever growing abundance of information became available to a rapidly growing population through the internet. It was no longer necessary to be tech savvy or to know what a relational database was in order to tap into this abundance: Millions of people started searching and connecting with friends on the web on a regular basis. This revolution is still causing significant social and economic changes, while evolving on two main tracks. The first track is the increase in storage, networking and processing power, harnessed in various physical configurations (data centers, handheld devices, personal computers and internet providers), making the same massive amounts of data appear virtually present simultaneously everywhere. The second force, complementing the first, is the algorithmic and statistical tools bridge between the people and the data.It is widely understood today that the survival of popular online applications (search engines, social networks, ad networks) heavily depends on the ability to represent data (e.g. indexing, classification and storage) and serve it (e.g. as query results, ads and recommendations) efficiently and reliably. It is hence not surprising that huge efforts are invested in analysis of massive amounts of data and people. These efforts transcend traditional borders of computer science and reposition the field as an engine for emerging multidisciplinary research, very much the way bioinformatics emerged before.We are already witnessing the economic and social influence of the internet on our lives, and the goal of this research is to identify simple structures that emerge in this ecosystem (graphs, clusters, high dimensional spaces, ratings, discrete choice and preference) and apply algorithmic tools to them.

Оригинален текст от CORDIS (на английски).

Участници

  • TECHNION - ISRAEL INSTITUTE OF TECHNOLOGY · HaifaКоординаторИзраел

Връзки

Данни: CORDIS, © Европейски съюз