Abstract
We present a bootstrapping algorithm to create a semantic lexicon from a list of seed words and a corpus that was mined from the web. We exploit extraction patterns to bootstrap the lexicon and use collocation statistics to dynamically score new lexicon entries. Extraction patterns are subsequently scored by calculating
the conditional probability in relation to a non-related text corpus. We find that verbs that are highly domain related achieved the highest accuracy and collocation statistics affect the accuracy positively and negatively during the bootstrapping runs.
the conditional probability in relation to a non-related text corpus. We find that verbs that are highly domain related achieved the highest accuracy and collocation statistics affect the accuracy positively and negatively during the bootstrapping runs.
| Original language | English |
|---|---|
| Title of host publication | Proceedings of the 8th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management |
| Editors | A. Fred, Jan Dietz, David Aveiro, Kecheng Liu, Jorge Bernardino, Joaquim Filipe |
| Place of Publication | Porto |
| Publisher | SciTePress |
| Pages | 189–196 |
| Volume | 1 |
| ISBN (Print) | 978-989-758-203-5 |
| DOIs | |
| Publication status | Published - 2016 |
Keywords
- Semantic Lexicon
- Bootstrapping
- Extraction Patterns
- Web Mining
Fingerprint
Dive into the research topics of 'Bootstrapping a Semantic Lexicon on Verb Similarities'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver