H2020Individual fellowship2017–2019

CoupledDB · High-Performance Indexing for Emerging GPU-Coupled Databases

Horizon 2020 — Marie Skłodowska-Curie Actions

Duration
2017-05-01 → 2019-04-30
EU contribution
€208,400
Participants
1
Scheme
MSCA-IF-EF-ST

Lines connect the coordinator with its partners.

Results in brief

High-Performance Indexing for Emerging GPU-Coupled Databases

"An unflagging trend over the past and probably next decade is the cheap availability of more and more random access memory (RAM). This is a pivotal difference to even a decade before, as it enables the computation *in memory* of expansive datasets, even on the scale of terabytes. Moreover, when computation can be done in memory, there are many more opportunities to exploit diverse processors concurrently to accelerate computation; suddenly, one can leverage highly parallel accelerators such as an Intel Xeon Phi or general purpose graphics processing cards (GPGPU's) and the inherent parallelism already present in any modern (i.e., super-scalar, vector-register-equipped, multi-core) computer. There are (at least) two *very* good reasons to target all of these opportunities for parallelism at once: (a) targetting diverse parallel devices produces more generic algorithms, data structures, and principles which can apply even to as-yet-uninvented parallel devices; and (b) one can use the *entire* computer to accelerate computation, rather than just heavily optimising one component while the others idle. At the same time, the ubiquity of mobile devices has generated an explosion of spatio-temporally-annotated textual content, such as geo-located tweets, ""check ins"" at local establishments, or news articles with spatial context. Supporting complex analytics over so much complex, heterogeneous data requires sophisticated data structures to help filter information *and* the use of parallelism to manage scale. Thus this project. The overall objective is to design spatio-temporal-textual indexes that are simultaneously architecture-conscious; i.e., can *fully* exploit the underlying hardware for maximal throughput---and suitable for multiple data-parallel platforms; i.e., can be ported, transferred, and simultaneously used by multiple components of the compute ecosystem. Such data structures would achieve the advantages of targetting multiple parallel devices and help manage the explosive growth in spatio-temporal-textual content."

Data: CORDIS, © European Union

Project objective

Index structures are foundational to the performance of database systems and large-scale simulations. Even small advances in indexing can therefore have widespread, sweeping impact on both industry competitiveness and scientific productivity. The confluence of several hardware trends is setting the stage for disruptive innovation in database indexing: deescalating costs of memory make it feasible to organise most of the ""hot"", frequently accessed data in memory rather than on disk; and increasingly commonplace accelerators such as graphics processing units (GPUs) offer large-scale parallelism with a lower energy footprint. Thus, in-memory indexing that exploits GPUs could be much cheaper, faster, and greener.However, effectively incorporating GPUs into computation is a principal research challenge. To idle the powerful multicore system in favour of exclusively using the GPU connected to it, as done currently, is to squander valuable resources. On the other hand, the GPU has a vastly different computational model, so cannot straight-forwardly leverage multicore techniques. The challenges in handling this dichotomy, in fact, will cross-cut many research areas as the heterogeneity in the compute ecosystem becomes ubiquitous in parallel processing.Building on preliminary results that suggest common data structures processed by architecture-specific algorithms can support heterogeneity, this action will design indexes for the coupled multicore-GPU database systems that will soon be ubiquitous. The indexes will enable more responsive simulations of complex objects such as neurons and vehicle trajectories and support the recent proliferation of mobile-generated data. Moreover, through the action, the researcher will transfer technical parallel programming skills to the host, while the host will transfer expertise about new data types to the researcher. The project results will contribute to Europe's positioning at the forefront of heterogeneous parallel processing.""

Original text from CORDIS.

Participants

  • NORGES TEKNISK-NATURVITENSKAPELIGE UNIVERSITET NTNU · TrondheimCoordinatorNorway

Links

Data: CORDIS, © European Union