GI LogoGI Logo
  • Anmelden
Digitale Bibliothek
    • Gesamter Bestand

      • Bereiche & Sammlungen
      • Titel
      • Autor
      • Erscheinungsdatum
      • Schlagwort
    • Diese Sammlung

      • Titel
      • Autor
      • Erscheinungsdatum
      • Schlagwort
Digital Bibliothek der Gesellschaft für Informatik e.V.
GI-DL
    • English
    • Deutsch
  • Deutsch 
    • English
    • Deutsch
Dokumentanzeige 
  •   Startseite
  • Fachbereiche
  • Datenbanken und Informationssysteme (DBIS)
  • Datenbank Spektrum
  • Datenbank Spektrum 19(2) - Juli 2019
  • Dokumentanzeige
JavaScript is disabled for your browser. Some features of this site may not work without it.
  •   Startseite
  • Fachbereiche
  • Datenbanken und Informationssysteme (DBIS)
  • Datenbank Spektrum
  • Datenbank Spektrum 19(2) - Juli 2019
  • Dokumentanzeige

Using the Semantic Web as a Source of Training Data

Autor(en):
Bizer, Christian [DBLP] ;
Primpeli, Anna [DBLP] ;
Peeters, Ralph [DBLP]
Zusammenfassung
Deep neural networks are increasingly used for tasks such as entity resolution, sentiment analysis, and information extraction. As the methods are rather training data hungry, it is necessary to use large training sets in order to enable the methods to play their strengths. Millions of websites have started to annotate structured data within HTML pages using the schema.org vocabulary. Popular types of entities that are annotated are products, reviews, events, people, hotels, and other local businesses [ 12 ]. These semantic annotations are used by all major search engines to display rich snippets in search results. This is also the main driver behind the wide-scale adoption of the annotation techniques. This article explores the potential of using semantic annotations from large numbers of websites as training data for supervised entity resolution, sentiment analysis, and information extraction methods. After giving an overview of the types of structured data that are available on the Semantic Web, we focus on the task of product matching in e‑commerce and explain how semantic annotations can be used to gather a large training dataset for product matching. The dataset consists of more than 20 million pairs of offers referring to the same products. The offers were extracted from 43 thousand e‑shops, that provide schema.org annotations including some form of product identifiers, such as manufacturer part numbers (MPNs), global trade item numbers (GTINs), or stock keeping units (SKUs). The dataset, which we offer for public download, is orders of magnitude larger than the Walmart-Amazon [ 7 ], Amazon-Google [ 10 ], and Abt-Buy [ 10 ] datasets that are widely used to evaluate product matching methods. We verify the utility of the dataset as training data by using it to replicate the recent result of Mudgal et al. [ 15 ] stating that embeddings and RNNs outperform traditional symbolic matching methods on tasks involving less structured data. After the case study on product data matching, we focus on sentiment analysis and information extraction and discuss how semantic annotations from the Web can be used as training data within both tasks.
  • Vollständige Referenz
  • BibTeX
Bizer, C., Primpeli, A. & Peeters, R., (2019). Using the Semantic Web as a Source of Training Data.   Datenbank-Spektrum: Vol. 19, No. 2. Springer. (S. 127-135). DOI: 10.1007/s13222-019-00313-y
@article{mci/Bizer2019,
author = {Bizer, Christian AND Primpeli, Anna AND Peeters, Ralph},
title = {Using the Semantic Web as a Source of Training Data},
journal = {Datenbank-Spektrum},
volume = {19},
number = {2},
year = {2019},
,
pages = { 127-135 } ,
doi = { 10.1007/s13222-019-00313-y }
}

Sollte hier kein Volltext (PDF) verlinkt sein, dann kann es sein, dass dieser aus verschiedenen Gruenden (z.B. Lizenzen oder Copyright) nur in einer anderen Digital Library verfuegbar ist. Versuchen Sie in diesem Fall einen Zugriff ueber die verlinkte DOI: 10.1007/s13222-019-00313-y

Haben Sie fehlerhafte Angaben entdeckt? Sagen Sie uns Bescheid: Feedback abschicken

Mehr Information

DOI: 10.1007/s13222-019-00313-y
ISSN: 1610-1995
Datum: 2019
Typ: Text/Journal Article

Keywords

  • Entity resolution
  • Information extraction
  • Product matching
  • Schema.org annotations
  • Semantic Web
  • Sentiment analysis
Sammlungen
  • Datenbank Spektrum 19(2) - Juli 2019 [10]

Zur Langanzeige


Über uns | FAQ | Hilfe | Impressum | Datenschutz

Gesellschaft für Informatik e.V. (GI), Kontakt: Geschäftsstelle der GI
Diese Digital Library basiert auf DSpace.

 

 


Über uns | FAQ | Hilfe | Impressum | Datenschutz

Gesellschaft für Informatik e.V. (GI), Kontakt: Geschäftsstelle der GI
Diese Digital Library basiert auf DSpace.