A Hybrid Information Extraction Approach Exploiting Structured Data Within a Text Mining Process
Zusammenfassung
Many data sets encompass structured data fields with embedded free text fields. The text fields allow customers and workers to input information which cannot be encoded in structured fields. Several approaches use structured and unstructured data in isolated analyses. The result of isolated mining of structured data fields misses crucial information encoded in free text. The result of isolated text mining often mainly repeats information already available from structured data. The actual information gain of isolated text mining is thus limited. The main drawback of both isolated approaches is that they may miss crucial information. The hybrid information extraction approach suggested in this paper adresses this issue. Instead of extracting information that in large parts was already available beforehand, it extracts new, valuable information from free texts. Our solution exploits results of analyzing structured data within the text mining process, i.e., structured information guides and improves the information extraction process on textual data. Our main contributions comprise the description of the concept of hybrid information extraction as well as a prototypical implementation and an evaluation with two real-world data sets from aftersales and production with English and German free text fields.
- Vollständige Referenz
- BibTeX
Kiefer, C., Reimann, P. & Mitschang, B.,
(2019).
A Hybrid Information Extraction Approach Exploiting Structured Data Within a Text Mining Process.
In:
Grust, T., Naumann, F., Böhm, A., Lehner, W., Härder, T., Rahm, E., Heuer, A., Klettke, M. & Meyer, H.
(Hrsg.),
BTW 2019.
Gesellschaft für Informatik, Bonn.
(S. 149-168).
DOI: 10.18420/btw2019-10
@inproceedings{mci/Kiefer2019,
author = {Kiefer, Cornelia AND Reimann, Peter AND Mitschang, Bernhard},
title = {A Hybrid Information Extraction Approach Exploiting Structured Data Within a Text Mining Process},
booktitle = {BTW 2019},
year = {2019},
editor = {Grust, Torsten AND Naumann, Felix AND Böhm, Alexander AND Lehner, Wolfgang AND Härder, Theo AND Rahm, Erhard AND Heuer, Andreas AND Klettke, Meike AND Meyer, Holger} ,
pages = { 149-168 } ,
doi = { 10.18420/btw2019-10 },
publisher = {Gesellschaft für Informatik, Bonn},
address = {}
}
author = {Kiefer, Cornelia AND Reimann, Peter AND Mitschang, Bernhard},
title = {A Hybrid Information Extraction Approach Exploiting Structured Data Within a Text Mining Process},
booktitle = {BTW 2019},
year = {2019},
editor = {Grust, Torsten AND Naumann, Felix AND Böhm, Alexander AND Lehner, Wolfgang AND Härder, Theo AND Rahm, Erhard AND Heuer, Andreas AND Klettke, Meike AND Meyer, Holger} ,
pages = { 149-168 } ,
doi = { 10.18420/btw2019-10 },
publisher = {Gesellschaft für Informatik, Bonn},
address = {}
}
Sollte hier kein Volltext (PDF) verlinkt sein, dann kann es sein, dass dieser aus verschiedenen Gruenden (z.B. Lizenzen oder Copyright) nur in einer anderen Digital Library verfuegbar ist. Versuchen Sie in diesem Fall einen Zugriff ueber die verlinkte DOI: 10.18420/btw2019-10
Haben Sie fehlerhafte Angaben entdeckt? Sagen Sie uns Bescheid: Feedback abschicken


(en)