Abstract
Tagged corpora are essential for evaluating and training natural language processing tools. The cost of constructing large enough manually tagged corpora is high, even when the annotation level is shallow. This article describes a simple method to automatically create a partially tagged corpus, using Wikipedia hyperlinks. The resulting corpus contains information about the correct segmentation of 523,599 non-consecutive words in 363,090 sentences. We used our method to construct a corpus of Modern Hebrew (which we have made available at http://www.cs.bgu.ac.il/-nlpproj). The method can also be applied to other languages where word segmentation is difficult to determine, such as East and South-East Asian languages.
| Original language | English |
|---|---|
| Title of host publication | Wikipedia and Artificial Intelligence |
| Subtitle of host publication | An Evolving Synergy - Papers from the 2008 AAAI Workshop |
| Pages | 61-63 |
| Number of pages | 3 |
| State | Published - 1 Dec 2008 |
| Event | 2008 AAAI Workshop - Chicago, IL, United States Duration: 13 Jul 2008 → 13 Jul 2008 |
Publication series
| Name | AAAI Workshop - Technical Report |
|---|---|
| Volume | WS-08-15 |
Conference
| Conference | 2008 AAAI Workshop |
|---|---|
| Country/Territory | United States |
| City | Chicago, IL |
| Period | 13/07/08 → 13/07/08 |
ASJC Scopus subject areas
- General Engineering
Fingerprint
Dive into the research topics of 'Using wikipedia links to construct word segmentation corpora'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver