GI LogoGI Logo
  • Anmelden
Digitale Bibliothek
    • Gesamter Bestand

      • Bereiche & Sammlungen
      • Titel
      • Autor
      • Erscheinungsdatum
      • Schlagwort
    • Diese Sammlung

      • Titel
      • Autor
      • Erscheinungsdatum
      • Schlagwort
Digital Bibliothek der Gesellschaft für Informatik e.V.
GI-DL
    • English
    • Deutsch
  • Deutsch 
    • English
    • Deutsch
Dokumentanzeige 
  •   Startseite
  • Lecture Notes in Informatics
  • Proceedings
  • BTW - Datenbanksysteme für Business, Technologie und Web
  • P331 - BTW2023- Datenbanksysteme für Business, Technologie und Web
  • Dokumentanzeige
JavaScript is disabled for your browser. Some features of this site may not work without it.
  •   Startseite
  • Lecture Notes in Informatics
  • Proceedings
  • BTW - Datenbanksysteme für Business, Technologie und Web
  • P331 - BTW2023- Datenbanksysteme für Business, Technologie und Web
  • Dokumentanzeige

Duplicate Table Discovery with Xash

Autor(en):
Koch, Maximilian [DBLP] ;
Esmailoghli, Mahdi [DBLP] ;
Auer, Sören [DBLP] ;
Abedjan, Ziawasch [DBLP]
Zusammenfassung
Data lakes are typically lightly curated and as such prone to data quality problems and inconsistencies. In particular, duplicate tables are common in most repositories. The goal of duplicate table detection is to identify those tables that display the same data.Comparing tables is generally quite expensive as the order of rows and columns might differ for otherwise identical tables. In this paper, we explore the application of Xash, a hash function previously proposed for the discovery of multi-column join candidates, for the use case of duplicate table detection. With Xash, it is possible to generate a so-called super key, which serves like a bloom filter and instantly identifies the existence of particular cell values. We show that using Xash it is possible to speed up the duplicate table detection process significantly. In comparison to other hash functions, such as SimHash and other competitors, Xash results in fewer false positive candidates.
  • Vollständige Referenz
  • BibTeX
Koch, M., Esmailoghli, M., Auer, S. & Abedjan, Z., (2023). Duplicate Table Discovery with Xash. In: König-Ries, B., Scherzinger, S., Lehner, W. & Vossen, G. (Hrsg.), BTW 2023. Gesellschaft für Informatik e.V.. DOI: 10.18420/BTW2023-18
@inproceedings{mci/Koch2023,
author = {Koch, Maximilian AND Esmailoghli, Mahdi AND Auer, Sören AND Abedjan, Ziawasch},
title = {Duplicate Table Discovery with Xash},
booktitle = {BTW 2023},
year = {2023},
editor = {König-Ries, Birgitta AND Scherzinger, Stefanie AND Lehner, Wolfgang AND Vossen, Gottfried} ,
doi = { 10.18420/BTW2023-18 },
publisher = {Gesellschaft für Informatik e.V.},
address = {}
}
DateienGroesseFormatAnzeige
B4-1.pdf590.9Kb PDF Öffnen

Sollte hier kein Volltext (PDF) verlinkt sein, dann kann es sein, dass dieser aus verschiedenen Gruenden (z.B. Lizenzen oder Copyright) nur in einer anderen Digital Library verfuegbar ist. Versuchen Sie in diesem Fall einen Zugriff ueber die verlinkte DOI: 10.18420/BTW2023-18

Haben Sie fehlerhafte Angaben entdeckt? Sagen Sie uns Bescheid: Feedback abschicken

Mehr Information

DOI: 10.18420/BTW2023-18
ISBN: 978-3-88579-725-8
Datum: 2023
Sprache: en (en)
Typ: Text/Conference Paper

Keywords

  • data discovery
  • data lakes
  • duplicate table detection
Sammlungen
  • P331 - BTW2023- Datenbanksysteme für Business, Technologie und Web [80]

Zur Langanzeige


Über uns | FAQ | Hilfe | Impressum | Datenschutz

Gesellschaft für Informatik e.V. (GI), Kontakt: Geschäftsstelle der GI
Diese Digital Library basiert auf DSpace.

 

 


Über uns | FAQ | Hilfe | Impressum | Datenschutz

Gesellschaft für Informatik e.V. (GI), Kontakt: Geschäftsstelle der GI
Diese Digital Library basiert auf DSpace.