Keyword extraction for text characterization
Zusammenfassung
Keywords are valuable means for characterizing texts. In order to extract keywords we propose an efficient and robust, languageand domainindependent approach which is based on small word parts (quadgrams). The basic algorithm can be improved by re-examining and re-ranking keywords using edit distance (i.e. Levenshtein distance) and an algorithm based on the relativistic addition of velocities (here: weights). For the purpose of evaluation, we compare our approach to frequency-based keyword extraction (exemplary text collection: 45000 intranet documents in German and English).
- Vollständige Referenz
- BibTeX
Renz, I., Ficzay, A. & Hitzler, H.,
(2003).
Keyword extraction for text characterization.
In:
Düsterhöft, A. & Thalheim, B.
(Hrsg.),
Natural language processing and information systems.
Bonn:
Gesellschaft für Informatik e.V..
(S. 228-234).
@inproceedings{mci/Renz2003,
author = {Renz, Ingrid AND Ficzay, Andrea AND Hitzler, Holger},
title = {Keyword extraction for text characterization},
booktitle = {Natural language processing and information systems},
year = {2003},
editor = {Düsterhöft, Antje AND Thalheim, Bernhard} ,
pages = { 228-234 },
publisher = {Gesellschaft für Informatik e.V.},
address = {Bonn}
}
author = {Renz, Ingrid AND Ficzay, Andrea AND Hitzler, Holger},
title = {Keyword extraction for text characterization},
booktitle = {Natural language processing and information systems},
year = {2003},
editor = {Düsterhöft, Antje AND Thalheim, Bernhard} ,
pages = { 228-234 },
publisher = {Gesellschaft für Informatik e.V.},
address = {Bonn}
}
| Dateien | Groesse | Format | Anzeige | |
|---|---|---|---|---|
| GI-Proceedings.29-19.pdf | 81.65Kb | Öffnen |
Haben Sie fehlerhafte Angaben entdeckt? Sagen Sie uns Bescheid: Feedback abschicken
Mehr Information
ISBN: 3-88579-358-X
ISSN: 1617-5468
Datum: 2003
Sprache:
(en)
(en)
Typ: Text/Conference Paper

