GI LogoGI Logo
  • Anmelden
Digitale Bibliothek
    • Gesamter Bestand

      • Bereiche & Sammlungen
      • Titel
      • Autor
      • Erscheinungsdatum
      • Schlagwort
    • Diese Sammlung

      • Titel
      • Autor
      • Erscheinungsdatum
      • Schlagwort
Digital Bibliothek der Gesellschaft für Informatik e.V.
GI-DL
    • English
    • Deutsch
  • Deutsch 
    • English
    • Deutsch
Dokumentanzeige 
  •   Startseite
  • Lecture Notes in Informatics
  • Proceedings
  • BTW - Datenbanksysteme für Business, Technologie und Web
  • P242 - BTW2015 - Datenbanksysteme für Business, Technologie und Web - Workshopband
  • Dokumentanzeige
JavaScript is disabled for your browser. Some features of this site may not work without it.
  •   Startseite
  • Lecture Notes in Informatics
  • Proceedings
  • BTW - Datenbanksysteme für Business, Technologie und Web
  • P242 - BTW2015 - Datenbanksysteme für Business, Technologie und Web - Workshopband
  • Dokumentanzeige

Ddup - towards a deduplication framework utilising apache spark

Autor(en):
Wilcke, Niklas [DBLP]
Zusammenfassung
This paper is about a new framework called DeduPlication (DduP). DduP aims to solve large scale deduplication problems on arbitrary data tuples. DduP tries to bridge the gap between big data, high performance and duplicate detection. At the moment a first prototype exists but the overall project status is work in progress. DduP utilises the promising successor of Apache Hadoop MapReduce [Had14], the Apache Spark Framework [ZCF+10] and its modules MLlib [MLl14] and GraphX [XCD+14]. The three main goals of this project are creating a prototype of the mentioned framework DduP, analysing the deduplication process about scalability and performance and evaluate the behaviour of different small cluster configurations. Tags: Duplicate Detection, Deduplication, Record Linkage, Machine Learning, Big Data, Apache Spark, MLlib, Scala, Hadoop, In-Memory
  • Vollständige Referenz
  • BibTeX
Wilcke, N., (2015). Ddup - towards a deduplication framework utilising apache spark. In: Ritter, N., Henrich, A., Lehner, W., Thor, A., Friedrich, S. & Wingerath, W. (Hrsg.), Datenbanksysteme für Business, Technologie und Web (BTW 2015) - Workshopband. Bonn: Gesellschaft für Informatik e.V.. (S. 253-262).
@inproceedings{mci/Wilcke2015,
author = {Wilcke, Niklas},
title = {Ddup - towards a deduplication framework utilising apache spark},
booktitle = {Datenbanksysteme für Business, Technologie und Web (BTW 2015) - Workshopband},
year = {2015},
editor = {Ritter, Norbert AND Henrich, Andreas AND Lehner, Wolfgang AND Thor, Andreas AND Friedrich, Steffen AND Wingerath, Wolfram} ,
pages = { 253-262 },
publisher = {Gesellschaft für Informatik e.V.},
address = {Bonn}
}
DateienGroesseFormatAnzeige
253.pdf96.18Kb PDF Öffnen

Haben Sie fehlerhafte Angaben entdeckt? Sagen Sie uns Bescheid: Feedback abschicken

Mehr Information

ISBN: 978-3-88579-636-7
ISSN: 1617-5468
Datum: 2015
Sprache: en (en)
Typ: Text/Conference Paper
Sammlungen
  • P242 - BTW2015 - Datenbanksysteme für Business, Technologie und Web - Workshopband [36]

Zur Langanzeige


Über uns | FAQ | Hilfe | Impressum | Datenschutz

Gesellschaft für Informatik e.V. (GI), Kontakt: Geschäftsstelle der GI
Diese Digital Library basiert auf DSpace.

 

 


Über uns | FAQ | Hilfe | Impressum | Datenschutz

Gesellschaft für Informatik e.V. (GI), Kontakt: Geschäftsstelle der GI
Diese Digital Library basiert auf DSpace.