PROC-LXML · Processing Large XML Data Sets: Algorithms and Limitations
6РП — Действия „Мария Кюри“
- Период
- 2005-08-31 → 2007-08-30
- Финансиране от ЕС
- 80 000 €
- Участници
- 1
- Схема
- IRG
Линиите свързват координатора с партньорите.
Накратко на български
Методите за търсене в големи масиви от данни, като например откриването на почти идентични уеб страници само по техните адреси. Това помага за по-точното измерване на качеството на търсачките и по-ефективното класиране на потребители в социалните мрежи.
Кратко обяснение, генерирано от езиков модел по текста на CORDIS. Оригиналът е по-долу.
Резултати накратко
Final Activity Report Summary - PROC-LXML (Processing Large XML Data Sets: Algorithms and Limitations)
In the PROC-LXML project we studied methods for searching large collections of web and XML documents. The main contributions of the project were the following: Search engine measurement. We developed a novel technique for estimating statistical parameters of a search engine by sending random queries to the search engine and analysing their results. This technique enables measurement and benchmarking of search engines without having to rely on their cooperation. We used the new technique to estimate the size, the freshness, and the quality of major search engines. One of the papers we published about this technique won the Best Paper Award at the International World-Wide Web Conference. Complexity of searching XML documents. We proved theoretical lower bounds on the amount of memory required to support searching over XML documents, which is significantly harder than searching over regular text documents. Ranking in social networks. We presented two ranking algorithms for social networks: one that ranks individuals based on their degree of influence in the network, and another that ranks groups, based on how cohesive and tightly knit these groups are. A paper about the latter algorithm won an honourable mention for the Best Application Paper Award at the International Conference on Data Mining. Detection of near-duplicate documents. We developed a highly efficient algorithm that detects near-duplicate web documents by examining only their URLs and without having to inspect their content.
Текст от CORDIS, на английски · Данни: CORDIS, © Европейски съюз
Цел на проекта
During the past few years, XML has become the dominant format for storing and exchanging information on the Internet. XML is often used to represent large text data sets, such as scientific corpora, repositories of Web pages, or streams of stock quotes. Processing large XML data sets efficiently has thus become one of the major challenges that researchers at the database, information retrieval, and WWW communities face today.This proposal focuses on three issues at the forefront of the XML research at the database community:1) Evaluation of queries over XML streams.2) Evaluation of queries over indexed XML data sets.3) Fast approximate evaluation of queries over XML data sets.The goals of the proposed project are three fold:1) Develop the first theoretical and systematic framework of lower bounds on the amount of resources needed to accomplish the above tasks.2) Exploit insights gained from the theoretical study to design more efficient and comprehensive algorithms that solve the above problems.3) Build an experimental system to test the proposed algorithms on real and artificial data.During the course of working on the project, I plan to continue existing collaborations in the area with researchers from the IBM Research Centre in California as well as to bring along new collaborators from the Technion whose areas of interest overlap the subject of the project. I plan to leverage on the expertise of my colleagues at the Technion in the areas of communication complexity, database, and information theory in order to obtain high quality results in this project.
Оригинален текст от CORDIS (на английски).
Участници
- TECHNION - ISRAEL INSTITUTE OF TECHNOLOGY · HAIFAКоординаторИзраел
Връзки
Данни: CORDIS, © Европейски съюз
