Automatic Diacritic Restoration for Resource-Scarce Languages

dc.contributor.author	De Pauw, G
dc.contributor.author	Wagacha, PW
dc.contributor.author	de Schryver, Gilles-Maurice
dc.date.accessioned	2013-06-21T09:17:11Z
dc.date.available	2013-06-21T09:17:11Z
dc.date.issued	2007
dc.identifier.citation	Lecture Notes in Computer Science Volume 4629, 2007, pp 170-179	en
dc.identifier.uri	http://link.springer.com/chapter/10.1007/978-3-540-74628-7_24
dc.identifier.uri	http://erepository.uonbi.ac.ke:8080/xmlui/handle/123456789/37318
dc.description.abstract	The orthography of many resource-scarce languages includes diacritically marked characters. Falling outside the scope of the standard Latin encoding, these characters are often represented in digital language resources as their unmarked equivalents. This renders corpus compilation more difficult, as these languages typically do not have the benefit of large electronic dictionaries to perform diacritic restoration. This paper describes experiments with a machine learning approach that is able to automatically restore diacritics on the basis of local graphemic context. We apply the method to the African languages of Cilubà, Gĩkũyũ, Kĩkamba, Maa, Sesotho sa Leboa, Tshivenda and Yoruba and contrast it with experiments on Czech, Dutch, French, German and Romanian, as well as Vietnamese and Chinese Pinyin.	en
dc.language.iso	en	en
dc.title	Automatic Diacritic Restoration for Resource-Scarce Languages	en
dc.type	Article	en