BNPNet · Bayesian nonparametric methods for networks and recommender systems
7РП — „Хора“ (Действия „Мария Кюри“)
- Период
- 2013-09-01 → 2015-08-31
- Финансиране от ЕС
- 231 283 €
- Участници
- 1
- Схема
- MC-IEF
Линиите свързват координатора с партньорите.
Накратко на български
Статистическите модели на сложни мрежи анализират връзките между обекти, като например приятелствата в социалните мрежи или взаимодействията между протеини. Те помагат за откриване на скрити групи и подобряване на точността при препоръчване на съдържание или идентифициране на липсващи връзки.
Кратко обяснение, генерирано от езиков модел по текста на CORDIS. Оригиналът е по-долу.
Резултати накратко
Bayesian nonparametric methods for networks and recommender systems
The last few years have seen a tremendous interest in the study and understanding of complex networks. Networks consist in a set of items, called vertices, with connections between them, called edges. Examples include the World Wide Web, social networks of friendship or other connections between individuals, organisational networks, networks of citations between papers, networks of shared interests in books or music, metabolic networks, food webs, etc. Statistical modeling of structured networks aims at building probabilistic network models characterized by a set of unknown parameters. These parameters may e.g. represent the popularity of a given node or its probability to be part of some community. Once inferred from the observed network, these parameters help to gain more insights on the structure of the network. It is of major importance for a broad number of applications ranging from social sciences to biology or information engineering. In social sciences, people are often interested in identifying communities from an observed network, such as Facebook or LinkedIn. In communication networks such as cell phone networks, network analysis may be used to detect latent terrorist cells. In computational biology, it can be used to find hidden groups in the protein-protein interaction network. This is valuable to gain better understanding of biological processes. Models can also be used to combine networks from heterogeneous data sources to improve the accuracy of predicted genetic interaction. Network models are often used to predict missing edges/nodes in the network. Applications include predicting missing edges in a business or a terrorist network. They are also particularly useful to build recommender systems which aim at providing automated targeted recommendations to individuals based on items they like. Recommender systems, which have gained a lot of attention over the past few years thanks to the famous Netflix prize, are now ubiquitous and used by companies like Amazon, Apple or Youtube. Recommender systems are of particular economic interest in business. According to the developer of the Amazon recommendation system, which pioneered automatic recommendations, 20% of the sales of Amazon in 2002 came from personalized recommendations. In this case, given a bipartite network of customers and items representing the purchase of a given item by a given customer, we aim at providing relevant recommendations to potential buyers of a given product. One may also be interested in obtaining a market segmentation of the customers and/or products, and in identifying trends in the evolution of the popularity of products. Realistic data models must be able to capture the salient features of real world graphs. Networks are often characterized in terms of their degree distribution. The degree of a given vertex is the number of connections of this vertex. Degree distributions of real world networks such as the Internet, telephone networks or co-authorship networks are often strongly non-Poissonian with a power-law behavior. In particular for customers/items networks, the distribution of the number of purchases of items is often heavy-tailed with a power-law behavior: most purchases only concern a small number of popular items. This is for example the case for book or music sales. Providing models that can adequately capture this power-law behavior is economically important, as useful information can be extracted from the tails of the distribution. One may be interested in identifying products with rising popularity, or niche markets with high potential. Moreover, networks typically have a very large number of vertices, and the number of edges scales linearly with the number of vertices. The degree of a given node is therefore much lower than the number of vertices, which can be considered infiinite comparatively. As an example, the Book-Crossing dataset, collected from the Book-Crossing community, contains about 1 million ratings of 278 000 users on 271 000 books, so on average less than 4 books per user. In e-commerce, the number of articles purchased by a single user is much smaller than the number of items available. We therefore need models that are scalable, and whose limit when the number of vertices becomes infinite still has remarkable statistical properties. In this project, we have developed novel classes of statistical models for networks than can (a) exhibit the power-law behavior of real-world networks, (b) have interpretable parameters, that can help to get a better understanding of the structure of the network, (c) whose large-scale properties are well understood, and (d) for which scalable algorithms for learning the parameters of the models are available. The main result is that the associated class of models can produce graphs which are called "sparse", where the number of connections in the network is much smaller than the maximum number of possible connections, which is what is expected from many real-world graphs. Models with this property have been proposed to deal with graphs with latent structure and dynamic graphs as well, but without associated algorithms for learning their parameters from data. We have shown that the models are able to learn and evaluate the level of sparsity of a range of real-world graphs, including social networks, biological networks or internet networks and to make more accurate predictions for graphs with these characteristics.
Текст от CORDIS, на английски · Данни: CORDIS, © Европейски съюз
Цел на проекта
Bayesian nonparametric (BNP) methods have become very popular over recent years in machine learning and statistics as it allows to build elegant and sophisticated models. Contrary to Bayesian parametric methods, this set of techniques allows the number of parameters to grow with the number of data and is particularly suitable in the data rich environment we now face. This project aims at developing new Bayesian models for the probabilistic modeling of large and structured data such as networks and buyer preferences.First, we aim at developing new models for networked data. The last few years have seen a tremendous interest in the study and understanding of complex networks. We plan to develop new models for static and dynamic networks, with or without clustering structure, that can handle a potentially large number of nodes and exhibit a power-law behavior, with simple inference procedures for the parameters.Second, we aim at developing BNP recommender systems. Recommender systems aim at predicting the preference that a user would give to a specific item. They are especially useful for e-commerce in order to provide targeted advertisements to users. When the number of potential users and items is potentially large compared to the number of transactions, a BNP approach becomes sensible. We aim at developing new probabilistic models for the modeling of the behavior of buyers over time.
Оригинален текст от CORDIS (на английски).
Участници
- THE CHANCELLOR, MASTERS AND SCHOLARS OF THE UNIVERSITY OF OXFORD · OxfordКоординаторОбединеното кралство
Връзки
Данни: CORDIS, © Европейски съюз
