Fact-Aware Document Retrieval for Information Extraction
Autor(en):
Zusammenfassung
Exploiting textual information from large document collections such as the Web with structured queries is an often requested, but still unsolved requirement of many users. We present BlueFact, a framework for efficiently retrieving documents containing structured, factual information from a full-text index. This is an essential building block for information extraction systems that enable ad-hoc analytical queries on unstructured text data as well as knowledge harvesting in a digital archive scenario.Our approach is based on the observation that documents share a set of common grammatical structures and words for expressing facts. Our system observes these keyword phrases using structural, syntactic, lexical and semantic features in an iterative, cost effective training process and systematically queries the search engine index with these automatically generated phrases. Next, BlueFact retrieves a list of document identifiers, combines observed keywords as evidence for a factual information and infers the relevance for each document identifier. Finally, we forward the documents in the order of their estimated relevance to an information extraction service. That way BlueFact can efficiently retrieve all the structured, factual information contained in an indexed collection of text documents.We report results of a comprehensive experimental evaluation over 20 different fact types on the Reuters News Corpus Volume I (RCV1). BlueFact’s scoring model and feature generation methods significantly outperform existing approaches in terms of fact retrieval performance. BlueFact fires significantly fewer queries against the index, requires significantly less execution time and achieves very high fact recall across different domains.
- Vollständige Referenz
- BibTeX
Boden, C., Löser, A., Nagel, C. & Pieper, S.,
(2012).
Fact-Aware Document Retrieval for Information Extraction.
Datenbank-Spektrum: Vol. 12, No. 2.
Springer.
(S. 89-100).
DOI: 10.1007/s13222-012-0088-4
@article{mci/Boden2012,
author = {Boden, Christoph AND Löser, Alexander AND Nagel, Christoph AND Pieper, Stephan},
title = {Fact-Aware Document Retrieval for Information Extraction},
journal = {Datenbank-Spektrum},
volume = {12},
number = {2},
year = {2012},
,
pages = { 89-100 } ,
doi = { 10.1007/s13222-012-0088-4 }
}
author = {Boden, Christoph AND Löser, Alexander AND Nagel, Christoph AND Pieper, Stephan},
title = {Fact-Aware Document Retrieval for Information Extraction},
journal = {Datenbank-Spektrum},
volume = {12},
number = {2},
year = {2012},
,
pages = { 89-100 } ,
doi = { 10.1007/s13222-012-0088-4 }
}
Sollte hier kein Volltext (PDF) verlinkt sein, dann kann es sein, dass dieser aus verschiedenen Gruenden (z.B. Lizenzen oder Copyright) nur in einer anderen Digital Library verfuegbar ist. Versuchen Sie in diesem Fall einen Zugriff ueber die verlinkte DOI: 10.1007/s13222-012-0088-4
Haben Sie fehlerhafte Angaben entdeckt? Sagen Sie uns Bescheid: Feedback abschicken
Mehr Information
ISSN: 1610-1995
Datum: 2012
Typ: Text/Journal Article

